DETAILED ACTION
This action is responsive to the claims filed on 12/12/2025. Claims 1-11, 13, and 15-19 are pending for examination.
Response to Arguments
Applicant argues:
“Ming… fails to read on the recited claim language in that it fails to evaluate a similarity vector to determine closeness of an output feature to the class prototypes. The other cited art cannot cure this deficiency.” (Remarks, page 9)
Applicant’s arguments have been fully considered and the examiner respectfully disagrees. Ming expressly computes similarity scores from distances to prototypes and forms a similarity vector. Specifically, Ming states: “To improve interpretability, we compute the similarity using: a_i = \exp(-d_i^2), which converts the distance to a score between 0 and 1,” (Ming, page 905, col. 2, paragraph 4) and further explains that this produces the “computed similarity vector a.” (Ming, page 905, col. 2, paragraph 6). Ming additionally uses that similarity vector directly in the model’s predictive operation (“With the computed similarity vector a = p(e), the fully connected layer computes z = Wa…” (Ming, page 905, col. 2, paragraph 5)).
Further, Ming teaches prototype-closeness regularization terms within its training objective that explicitly determine and enforce closeness between encoded instances (i.e., output features/embeddings) and prototypes. Ming describes a clustering regularization R_c that “minimiz[es] the squared distance between an encoded instance and its closest prototype,” and an evidence regularization R_e that “encourages each prototype vector to be as close to an encoded instance as possible.” (Ming, page 906, col. 1, paragraph 1). Ming further confirms these regularization terms are jointly minimized as part of the overall optimized loss function.
Accordingly, Ming teaches that prototype-closeness is determined during training using regularization losses operating on the relationships between encoded instances and prototypes, and Ming also expressly represents those relationships as a “similarity vector a” computed from the corresponding prototype distances. Thus, Ming’s optimized objective evaluates the similarity relationships (expressed as the similarity vector and/or its underlying distances) to determine and enforce closeness of the encoded output feature to the prototypes, as required by the amended limitation.
Applicant’s remarks regarding the provisional double patenting rejections have been fully considered. The terminal disclaimers filed to obviate the double patenting with respect to applications 18479326, 17158466, and 18479372 have been entered. Accordingly, the double patenting rejections are withdrawn.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-11, 13, and 15-19 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Statutory Categories
Claims 1-11, 13, and 15-19 are directed to a method.
Independent Claims – Claim 1
Step 2A Prong 1: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes. Independent claim 1 recites limitations that are abstract ideas in the form of mental processes:
Claim 1 recites:
Clustering… the plurality of input segments into clusters using respective resolution-controllable class prototypes allocated to each of a plurality of classes, each of the respective resolution-controllable class prototypes including a respective subset of the output features that characterizes a respective associated one of the plurality of classes; (clustering the input segments into classes stated a high level of generality on how to perform the clustering is being considered a mental process of evaluation (clustering) that can reasonably be performed by the human mind)
calculating, using the clusters, similarity scores that indicate a similarity of a given one of the output features to a given one of the respective resolution-controllable class prototypes responsive to distances, in a latent space, between the output feature and the respective resolution-controllable class prototypes; (calculating similarity scores using clusters based on numerical distances merely involves mathematical calculations, algorithms, or formulas. Paragraphs 40-43 of the specification provide the related mathematical disclosure.)
pushing together the plurality of segments to corresponding ones of the respective resolution-controllable class prototypes under a constraint that pushing is only to occur toward the corresponding ones of the respective-controllable class prototypes having a same class and further under at least one distance-based loss function (this limitation merely recites mathematical concepts in the form of mathematical formulas, algorithms, or calculations. Page 15 of the specification provides support for the mathematical disclosure.)
performing… a prediction and prediction support operation that provides a value of prediction and an interpretation for the value of prediction responsive to the input segments and the similarity score by solving an optimization problem having a cross-entropy loss (lc), a gradient loss (la), a closeness loss (lc) and a diversity regularization term (ld), wherein the interpretation for the value of prediction is provided using only non-negative weights and lacking a weight bias in the fully connected layer, wherein the closeness loss evaluates a similarity vector to determine closeness of the given one of the output features to the respective resolution-controllable class prototypes (this limitation merely recites mathematical concepts in the form of mathematical formulas, algorithms, or calculations. Pages 14-17 of the specification provides support for the mathematical disclosure.)
This claim further recites the following additional elements for the purposes of Step 2A Prong Two analysis:
converting, by a convolutional layer having one or more filters and a sliding window, an input data sequence having a plurality of input segments into a set of output features, the input data sequence being electric health records; (this limitation invokes convolutional layers merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
in multiple protype storage elements, (this limitation invokes protype storage elements merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
performing, by a fully connected layer (this limitation invokes fully connected layers merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
The additional limitations fail step 2A Prong 2 of the 101 analysis because they do not transform the claim into a practical application. These limitations are too abstract or lack technical improvement that would make the concept practically useful. Without clear utility or integration into a specific field, the claim does not relate to any particular application. It does not meet the requirements of Step 2A Prong 2, as it fails to make the concept meaningfully applicable in practice. Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is directed to an abstract idea.
This claim recites the following additional elements for the purposes of Step 2B analysis:
converting, by a convolutional layer having one or more filters and a sliding window, an input data sequence having a plurality of input segments into a set of output features, the input data sequence being electric health records; (this limitation invokes convolutional layers merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
in multiple protype storage elements, (this limitation invokes protype storage elements merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
performing, by a fully connected layer (this limitation invokes fully connected layers merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
The claim also fails Step 2B of the analysis because the additional limitations do not amount to significantly more than the abstract idea itself. The additional limitations do not enhance the claim in a way that would move it beyond its abstract ideas as they minimally elaborate on the core concept without adding any inventive or technical substance. Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Dependents of Claim 1
The remaining dependent claims corresponding to independent claim 1 do not recite additional elements, whether considered individually or in combination, that are sufficient to integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. The analysis of which is shown below:
The claims below recite additional limitations which fail step 2A Prong 2 of the 101 analysis because they do not transform the claim into a practical application. These limitations are too abstract or lack technical improvement that would make the concept practically useful. Without clear utility or integration into a specific field, the claim does not relate to any particular application. It does not meet the requirements of Step 2A Prong 2, as it fails to make the concept meaningfully applicable in practice.
The claims also fails Step 2B of the analysis because the additional limitations do not amount to significantly more than the abstract idea itself. The additional limitations do not enhance the claim in a way that would move it beyond its abstract ideas as they minimally elaborate on the core concept without adding any inventive or technical substance. The claims are unpatentable.
Claim 2 recites the further limitation of:
The computer-implemented method of claim 1, wherein the performing performs the prediction and prediction support operation that provides the value of prediction and an interpretation for the value of prediction responsive to the input segments and the similarity vector (this further limitation on an aforementioned abstract idea is still being considered as a mental process of evaluation which can reasonably be performed in one’s mind or with aid of pen and paper)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 3 recites the further limitation of:
The computer-implemented method of claim 2, wherein the performing performs,… the prediction and prediction support operation, wherein the interpretation for the value of prediction is provided using only non-negative weights and lacking a weight bias in the fully connected layer (this further limitation on an aforementioned abstract idea is still being considered as a mental process of evaluation which can reasonably be performed in one’s mind or with aid of pen and paper)
by a fully connected layer, (For the purposes of Step 2A Prong 2 and Step 2B: this limitation invokes fully connected layers merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 4 recites the further limitation of:
The computer-implemented method of claim 1, wherein the set of output features is represented by a non-linear function plus a bias term (this limitation merely recites mathematical concepts in the form of mathematical formulas, algorithms, or calculations. Pages 11-12 of the specification provides support for the mathematical disclosure)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 5 recites the further limitation of:
The computer-implemented method of claim 1, wherein each of the class prototypes collectively form a class prototype vector that is a latent representation of a prototypical segment learned through gradient descent (this limitation merely recites mathematical concepts in the form of mathematical formulas, algorithms, or calculations. Paragraphs 53-59 of the specification provide the related mathematical disclosure.)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 6 recites the further limitation of:
The computer-implemented method of claim 1, wherein each of the respective resolution-controllable class prototypes has a selectable resolution corresponding to an associated one of the one or more filters (this limitation invokes filters merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 7 recites the further limitation of:
The computer-implemented method of claim 1, wherein a dimensionality of each of the respective resolution-controllable class prototypes is equal to a dimensionality of each of the output features allocated thereto (For the purposes of Step 2A Prong 2 and Step 2B: this limitation invokes dimensionality merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 8 recites the further limitation of:
The computer-implemented method of claim 1, wherein a number of the multiple protype storage elements is equal to a number of the one or more filters in the convolutional layer (For the purposes of Step 2A Prong 2 and Step 2B: this limitation invokes storage elements and filters merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 9 recites the further limitation of:
The computer-implemented method of claim 1, wherein the one or more filters comprise multiple filters having different size lengths configured to selectively address different resolutions of the plurality of input segments corresponding to different levels of granularity (For the purposes of Step 2A Prong 2 and Step 2B: this limitation invokes different size filters merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 10 recites the further limitation of:
The computer-implemented method of claim 1, wherein the similarity scores range 0 to 1, wherein a 0 indicates that the given one of the output features is different from the given one of the respective resolution-controllable class prototypes, and a 1 indicates that the given one of the output features is identical to the given one of the respective resolution-controllable class prototypes. (this limitation merely recites mathematical concepts in the form of mathematical formulas, algorithms, or calculations. Page 13 of the specification provides support for the mathematical disclosure)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 11 recites the further limitation of:
The computer-implemented method of claim 2, further comprising performing a max pooling operation on the similarity cores to obtain the similarity vector (this limitation merely recites mathematical concepts in the form of mathematical formulas, algorithms, or calculations. Paragraph 41 of the specification provides support for the mathematical disclosure)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 13 recites the further limitation of:
The computer-implemented method of claim 3, further applying a softmax operation to an output of the fully connected layer to obtain the value of prediction and the interpretation for the value of prediction (this limitation merely recites mathematical concepts in the form of mathematical formulas, algorithms, or calculations. Paragraph 45 of the specification provides support for the mathematical disclosure)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 15 recites the further limitation of:
The computer-implemented method of claim 1, wherein said performing step is performed to solve an optimization problem having an accuracy component and an interpretability component (this limitation merely recites mathematical concepts in the form of mathematical formulas, algorithms, or calculations. Paragraph 47 of the specification provides support for the mathematical disclosure)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 16 recites the further limitation of:
The computer-implemented method of claim 1, wherein the diversity regularization term penalizes small distances between the respective resolution-controllable class prototypes below a diversity regularization threshold distance (this limitation merely recites mathematical concepts in the form of mathematical formulas, algorithms, or calculations. Paragraph 55 of the specification provides support for the mathematical disclosure)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 17 recites the further limitation of:
The computer-implemented method of claim 1, further comprising performing a control action responsive to the prediction and prediction support operation (For the purposes of Step 2A Prong 2 and Step 2B: this limitation invokes control actions merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 18 recites the further limitation of:
The computer-implemented method of claim 17, wherein the control action is selected from the group consisting of turning off an impending failing element, swapping out a failed component for another operating component, and switching to a secure network (this further limitation on an aforementioned abstract idea of a selection a control action and is still being considered as a mental process of evaluation which can reasonably be performed in one’s mind or with aid of pen and paper)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 19 recites the further limitation of:
The computer-implemented method of claim 1, wherein the convolutional layer has a plurality of filters of different sizes (this limitation invokes filters for a convolutional layer merely as a tool to perform an existing process and is considered as mere instructions to apply an exception, see MPEP 2106.05(f))
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 4, 6, and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Nguyen et al. ("Deepr : A Convolutional Net for Medical Records”), hereafter referred to as Nguyen, in view of Redmon et al. (US 20190102646 A1), hereafter referred to as Redmon, and in further view of Saralajew et al. (Saralajew, S., Holdijk, L., Rees, M., & Villmann, T. (2018). Prototype-based Neural Network Layers: Incorporating Vector Quantization. ArXiv, abs/1812.01214), hereafter referred to as Saralajew, and Ming et al. (Ming, Y., Xu, P., Qu, H., & Ren, L. (2019, July). Interpretable and steerable sequence learning via prototypes. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (pp. 903-913).), hereafter referred to as Ming.
Regarding claim 1, Nguyen teaches the following limitations:
converting, by a convolutional layer having one or more filters and a sliding window, the input data sequence into a set of the output features. (Nguyen, page 25, col. 1, section C2, “On top of the word embedding layers is a convolutional layer. Each convolution operation reads a sliding window of size 2d + 1 and produces p filter responses”, a convolutional layer is responsible for processing input of medical records with a sliding window.)
the input data sequence being electric health records (Nguyen, page 23, col. 2, section 3A, “Deepr is a multilayered architecture based on CNNs. The information flow is summarized in Fig. 1. At the bottom level, Deepr sequences the EMR into a “sentence,” or equivalently, a sequence of “words.”.”, input for the predictive model comprises EMR’s or Electronic Medical Records.)
clustering, in multiple prototype storage elements, a plurality of input segments into clusters using respective resolution-controllable class prototypes allocated to each of a plurality of classes, each of the respective resolution-controllable class prototypes including a respective subset of output features that characterizing a respective associated one of the plurality of classes; (Nguyen, page 6, col. 1, paragraph 1, “Deepr discovers disease clusters which partly correspond to nodes in the ICD10 hierarchy. Apart from pregnancy, child birth issues and injuries, the conditions are not totally separately suggesting a complex dependencies in the disease space.”, clusters are formed from EMR information to detect output features like certain diseases.)
Redmon, in the same field of clustering with convolutional neural networks, teaches the following limitations which Nguyen fails to teach:
calculating, using the clusters, similarity scores that indicate a similarity of a given one of the output features to a given one of the respective resolution-controllable class prototypes responsive to distances, in a latent space, between the output feature and the respective resolution-controllable class prototypes (Redmon, paragraph 78, “The process 600 includes determining 610 priors for a set of bounding boxes by performing a clustering analysis of bounding boxes in localization labels from a corpus of training images. The clustering analysis may use a distance metric based on intersection over union. Instead of choosing priors by hand, a clustering analysis (e.g., a k-means clustering analysis) can be run on the training set bounding boxes to automatically find good priors for the bounding boxes of the convolutional neural network. If standard k-means with Euclidean distance is used, then larger boxes generate more error than smaller boxes. However, having priors for the bounding boxes that lead to good intersection over union (IOU) scores, which is independent of the size of the box, may be advantageous. For example, the distance metric used in the clustering analysis”, based on the clusters, IOU scores based on Euclidean distance is calculated to indicate a similarity of a given output to a predetermined class prototype (localization label).
Redmon, paragraph 35, “Some classifier networks may be trained at 224×224, and the resolution may be increased to 448×448 for object detection. This means the network has to simultaneously switch to learning object detection and adjust to the new input resolution. In some implementations, the convolutional neural network 110 is first fine-tuned by training with images from a classification dataset (e.g., ImageNet) but operating at the full resolution used for detection (e.g., 448×448) for 10 epochs. This approach may give the convolutional neural network 110 time to adjust its filters to work better on higher resolution input. The resulting model stored in the convolutional neural network 110 by training may then be fine-tuned with images from detection datasets (e.g., COCO). For example, training the convolutional neural network 110 with images from an object detection dataset that has a classification label but lacks a localization label may include up-sampling the training image to a higher resolution to match a resolution of training images in a corpus of object detection training images that are associated with classification labels and localization labels.”, resolution is controlled via convolutional filters to adjust the resolution, suggesting that the class prototypes are resolution-controllable.);
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to have incorporated the teachings disclosed by Nguyen with the teachings disclosed by Redmon (i.e., clustering using protypes based on distance). A motivation for the combination is to provide better accuracy of estimated features by incorporating images of varying resolutions. (Redmon, paragraph 28, “For example, using multi-scale training, the same convolutional neural network model can be applied to images at varying sizes or resolutions, providing a smooth tradeoff between speed and accuracy.”)
Saralajew, in the same field of neural network classification, teaches the following limitations which Nguyen and Redmon fails to teach:
pushing the plurality of segments to corresponding ones of the respective resolution- controllable class prototypes under a constraint that pushing is only to occur toward the corresponding ones of the respective-controllable class prototypes having a same class and further under at least one distance-based loss function; (Saralajew, page 6, section 6.1, “At inference, the distance between the output of the feature extraction layer and all learned prototypes is used for classification. This requires the selection of a differentiable loss function and distance measure.”,
Saralajew, page 3, paragraph 2, “In LVQ, each prototype wk is additionally equipped with a class label ck ∈ C = {1,2,...,NC} and the prototypes are distributed regarding a training dataset X = {(x,c(x))|x ∈ Rn,c(x) ∈ C}. The class of a given data point x is defined as
PNG
media_image1.png
18
118
media_image1.png
Greyscale
”, Saralajew’s LVQ framework explicitly associates each prototype w_k with a class label c_k and defines classification of a data point x by the winner prototype κ(x) via
PNG
media_image1.png
18
118
media_image1.png
Greyscale
, which is a same-class association between an input segment and the corresponding class-labeled prototype. Saralajew further states that classification uses distances between the feature-extraction outputs and prototypes and requires a differentiable loss function and distance measure. Thus, training/optimization based on the distance/loss framework necessarily enforces movement (“pushing” under BRI) of the encoded segments with respect to the class-labeled prototypes.)
performing, by a fully connected layer a prediction… that provides a value of prediction…responsive to the input segments and the similarity score (Saralajew, page 6, section 6.1, “At inference, the distance between the output of the feature extraction layer and all learned prototypes is used for classification... In general, the output vector o(x) of (7) in the last FCL of a NN is element-wise normalized by the softmax activation… with ˆp(x) ∈ [0,1]NC being a probability vector of the estimated class probabilities”, Saralajew explicitly describe the prediction operation as producing a class-probability vector (i.e., a “value of prediction”) via the last fully connected layer output and softmax, and describe classification based on distances to learned prototypes, which corresponds to being responsive to the feature-extraction outputs and prototype distances.)
a gradient loss (la), (Saralajew, page 12, paragraph 5, “For some of the regularization and loss terms described in this report it is needed to estimate the data distribution over the training dataset. Since the network is optimized by stochastic gradient descent learning, the statistics over the whole dataset cannot be estimated during run-time. This can be compensated via moving averages/ moving variances. Additionally, we frequently prefer to perform a zero de-biasing according to [94] to avoid biased gradients.”,)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to have incorporated the teachings disclosed by Nguyen and Redmon with the teachings disclosed by Saralajew. A motivation for the combination is to improve the quality of decision boundaries gained from prototype vectors. (Saralajew, page 9, paragraph 6, “First experimental numerical results show that using this latter regularization approach forces the prototypes to be both generative and discriminative in the projection space at the same time.”)
Nguyen teaches a machine learning model that converts input sequences into a set of output features using a sliding window. Ming, in the same field of neural network classification, teaches the following limitations which the above prior art fails to teach:
performing… [a] prediction support operation that provides a value of prediction and an interpretation for the value of prediction responsive to the input segments and the similarity score (Ming, page 906, col. 1, last paragraph, “ProSeNet is readily explainable by consulting the most similar prototypes. When making predictions based on a new input sequence, the explanation can be generated along with the inference procedure. A prediction could be explained by a weighted addition of the contribution of the most similar prototypes: Input: Prediction: Explanation: pizza is good but service is extremely slow Negative 0.69 * good food but worst service (Negative 2.1) + 0.30 * service is really slow (Negative 1.1). The factors in front of the prototype sequences are the similarities between the input and the prototypes. At the end of each prototype shows its associated weights wi. The weights can be interpreted as the model’s confidence on the possible labels of the prototype.”, Ming expressly discloses a support/explanation operation accompanying inference that explains a prediction using similarities to prototypes and associated weights, i.e., an interpretation for the value of prediction that is responsive to the input and prototype similarity values.)
by solving an optimization problem having a cross-entropy loss (le) (Ming, page 905, section 3.2, “For accuracy, we minimize the cross-entropy loss on training set:
PNG
media_image2.png
25
356
media_image2.png
Greyscale
”
Page 906, col. 1, paragraph 3, “Full objective. To summarize, the loss that we are minimizing is:
PNG
media_image3.png
47
368
media_image3.png
Greyscale
”, Ming includes a cross-entropy loss (CE) as the primary accuracy term in their overall loss function, forming part of a unified objective optimized during training. The loss function being minimized contains a cross-entropy term along with others.)
a closeness loss (lc), (Ming, page 905, col. 2, paragraph 3, “To improve interpretability, we compute the similarity using:
PNG
media_image4.png
19
80
media_image4.png
Greyscale
which converts the distance to a score between 0 and 1.”, discloses how closeness loss is used by computing the distance between extracted features and learned prototypes.)
and a diversity regularization term (ld). (Ming, page 905, section 3.2, “We prevent such phenomenon through a diversity regularization term that penalizes on prototypes that are close to each other:
PNG
media_image5.png
57
306
media_image5.png
Greyscale
… Rd is a soft regularization that exerts a larger penalty on smaller pairwise distances. By keeping prototypes distributed in the latent space, it also helps produce a sparser similarity vector a.”, a “diversity regularization” term (Rd) which penalizes prototypes that are too close to each other in the latent space, thereby explicitly encouraging diversity among prototypes.)
wherein the interpretation for the value of prediction is provided using only non-negative weights (Ming, page 905, col. 2, section 3.2, paragraph 3, “Sparsity and non-negativity. In addition, to further enhance interpretability, we add L1 penalty on the fully connected layer f , and constrain the weight matrix W to be non-negative. The L1 sparsity penalty and non-negative constraints on f help to learn sequence prototypes that have more unitary and additive semantics for classification.”, Ming requires that the fully connected layer’s weights are to be non-negative.)
and lacking a weight bias in the fully connected layer (Ming, page 905, col. 2, paragraph 3, “With the computed similarity vector a = p(e), the fully connected layer computes z = Wa, where W is a C × k weight matrix and C is the output size (i.e., the number of classes in classification tasks). To enhance interpretability, we constrain W to be nonnegative. For multi-class classification tasks, a softmax layer is used to compute the predicted probability: yˆ i = exp(zi)/ÍC j=1 exp(zj).”, the fully connected layer is explicitly lacking a weight bias.).
wherein the closeness loss evaluates the similarity vector to determine closeness of the given one of the output features to the respective… class prototypes. (Ming, page 905, col. 2, paragraph 3, “To improve interpretability, we compute the similarity using:
PNG
media_image4.png
19
80
media_image4.png
Greyscale
which converts the distance to a score between 0 and 1. Zero can be interpreted as the sequence embedding e being completely different from the prototype vector pi , and one means they are identical. With the computed similarity vector a = p(e), the fully connected layer computes z = Wa, where W is a C × k weight matrix and C is the output size (i.e., the number of classes in classification tasks).”, Ming discloses that similarity between an encoded instance and each prototype is computed from the distance as:
PNG
media_image4.png
19
80
media_image4.png
Greyscale
Ming explicitly defines and utilizes a similarity vector representing similarity between encoded output features and respective prototypes.
Ming, page 906, col. 1, paragraph 3, “Full objective. To summarize, the loss that we are minimizing is:
PNG
media_image6.png
32
265
media_image6.png
Greyscale
”, Because Ming’s similarity vector is computed directly from the prototype distances (a_i = exp(−d_i^2)), and because the clustering and evidence regularization losses minimize those same distances within the unified optimization objective, Ming’s closeness regularization necessarily evaluates the similarity relationships embodied in the similarity vector to determine closeness between encoded features and prototypes. It is interpreted by the Examiner that Ming’s prototype-closeness regularization evaluates the similarity relationships represented in the similarity vector to determine and enforce closeness between encoded output features and respective prototypes. Because Ming’s similarity vector is computed directly from the prototype distances (a_i = exp(−d_i^2)), and because the clustering and evidence regularization losses minimize those same distances within the unified optimization objective, Ming’s closeness regularization necessarily evaluates the similarity relationships embodied in the similarity vector to determine closeness between encoded features and prototypes.)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings disclosed by Nguyen, Redmon, and Saralajew with the teachings disclosed by Ming (i.e., prediction interpretation with non-negative weights and no bias). Since Ming et al. design their loss function to be modular and open to additional regularization terms, it would have been obvious to one of ordinary skill in the art to add a gradient loss to their unified loss function. It should be noted that a proper obviousness rejection under 35 U.S.C. 103 does not require that all claimed features be expressly or bodily incorporated in a single reference or disclosed in precisely the same way as claimed. See MPEP § 2145 (III) “The test for obviousness is not whether the features of a secondary reference may be bodily incorporated into the structure of the primary reference.... Rather, the test is what the combined teachings of those references would have suggested to those of ordinary skill in the art.” Here, the combination of Nguyen with Ming demonstrates the routine and well-known practice of unifying multiple loss terms in a single optimization objective for neural network training, and the addition of further regularization or gradient-based loss terms would have been well within the ordinary skill in the art. A motivation for the combination is to provide increased interpretability of machine learning parameters. (Ming, page 906, col. 1, paragraph 1, “To improve interpretability, Li et al.[19] also proposed two regularization terms to be jointly minimized, the clustering regularization Rc and the evidence regularization Re . Rc encourages a clustering structure in the latent space by minimizing the squared distance between an encoded instance and its closest prototype”)
Regarding claim 4, Nguyen, Redmon, Saralajew, and Ming teaches the limitations of claim 1. Nguyen further teaches:
wherein the set of output features is represented by a non-linear function plus a bias term (Nguyen, page 3, col. 2, section C, part d, “Each convolution operation reads a
sliding window of size 2d + 1 and produces p filter responses
PNG
media_image7.png
60
202
media_image7.png
Greyscale
… b is bias”, it is interpreted by the examiner that filter responses in a convolutional operation are synonymous with output features, as they represent the results of applying the filter to the input data and form the feature map of a convolutional layer.)
Regarding claim 6, Nguyen, Redmon, Saralajew, and Ming teaches the limitations of claim 1. Redmon further teaches:
wherein each of the respective resolution-controllable class prototypes has a selectable resolution corresponding to an associated one of the one or more filters. (Redmon, paragraph 35, “Some classifier networks may be trained at 224×224, and the resolution may be increased to 448×448 for object detection. This means the network has to simultaneously switch to learning object detection and adjust to the new input resolution. In some implementations, the convolutional neural network 110 is first fine-tuned by training with images from a classification dataset (e.g., ImageNet) but operating at the full resolution used for detection (e.g., 448×448) for 10 epochs. This approach may give the convolutional neural network 110 time to adjust its filters to work better on higher resolution input. The resulting model stored in the convolutional neural network 110 by training may then be fine-tuned with images from detection datasets (e.g., COCO). For example, training the convolutional neural network 110 with images from an object detection dataset that has a classification label but lacks a localization label may include up-sampling the training image to a higher resolution to match a resolution of training images in a corpus of object detection training images that are associated with classification labels and localization labels.”, resolution is controlled via convolutional filters to adjust the resolution according to a corpus of multi-resolution training images.)
Regarding claim 8, Nguyen, Redmon, Saralajew, and Ming teaches the limitations of claim 1. Saralajew further teaches:
wherein a number of the multiple protype storage elements is equal to a number of the one or more filters in the convolutional layer (Saralajew, page 7, section 6.2, paragraph 3, “For a kernel-prototype convolution the number of filters Nf equals the number of kernel-prototypes NW.”).
Claims 7, 10, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Nguyen in view of Redmon, and in further view of Saralajew, and Ming, as applied to claims 1, 4, 6, and 8 above, and in further view of Pai et al. (US 11580420), hereafter referred to as Pai.
Regarding claim 7, Nguyen, Redmon, Saralajew, and Ming teaches the limitations of claim 1. Pai further teaches:
wherein a dimensionality of each of the respective resolution-controllable class prototypes is equal to a dimensionality of each of the output features allocated thereto (Pai, col. 6, lines 4-13, “As used herein, the term “feature space” refers to a space reflecting features of data points. Specifically, a feature space can reflect one or more dimensions corresponding to one or more features. For instance, the model analysis system can map features of data points to locations within different dimensions of a feature space. For instance, for a dataset include a set of ten features for each of the data points, a feature space can include ten dimensions to which each of the data points is mapped based on the values of the corresponding features.”, output features have dimensions based on their corresponding features (prototypes)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combined teachings of Nguyen, Redmon, Saralajew, and Ming with the teachings of Pai it is directed to selecting and using prototypes via an explicit optimization objective that balances (i) selecting prototypes that represent similar outcomes/labels and (ii) selecting the fewest prototypes. Pai expressly states that its objective function criteria ensure “data points are covered … by prototypes closer to them (in the label space)” (Pai, col. 12, lines 35-40) and penalize the number of prototypes so the system “selects the fewest possible number of prototypes,” (Pai, col. 11, lines 57-58) which teaches an optimization having both performance/accuracy-related and interpretability/efficiency-related components. A motivation for the combination would have been to improve interpretability while maintaining accurate representation of model behavior, consistent with Pai’s recognition of an “interpretability and accuracy” trade-off and its stated goal of providing analysis “without sacrificing either accuracy or interpretability.” (Pai, col. 5, line 31).
Regarding claim 10, Nguyen, Redmon, Saralajew, and Ming teaches the limitations of claim 1. Pai teaches:
wherein the similarity scores range from 0 to 1, (Pai, col. 9, lines 41-45, “Additionally, pre-processing the data points can include the model analysis system 102 scaling the target (e.g., outputs of the machine-learning model) to have values between 0 and 1 with ymax=1 and ymin=0.”, Pai explicitly discloses scaling model outputs/targets to be between 0 and 1, which teaches that the scores can be constrained to the 0–1 range.).
The rationale for combining Nguyen, Redmon, Saralajew, and Ming with Pai is similar to that as applied for claim 7 above.
Ming further teaches:
wherein a 0 indicates that the given one of the output features is different from the given one of the respective resolution-controllable class prototypes, and a 1 indicates that the given one of the output features is identical to the given one of the respective resolution-controllable class prototypes (Ming, page 905, col. 2, paragraph 3, “To improve interpretability, we compute the similarity using ai = exp(−d2 i ), which converts the distance to a score between 0 and 1. Zero can be interpreted as the sequence embedding e being completely different from the prototype vector pi , and one means they are identical.”, Ming expressly states the 0–1 score interpretation relative to prototypes: 0 = completely different, 1 = identical, matching the claim’s semantics.).
Regarding claim 15, Nguyen, Redmon, Saralajew, and Ming teaches the limitations of claim 1. Pai further teaches:
wherein said performing step is performed to solve an optimization problem having an accuracy component and an interpretability component (Pai, col. 4, lines 58-64, “Specifically, the model analysis system improves accuracy by using an iterative process to select, from a plurality of data points in a dataset, a plurality of representative prototypes based on distances between data points, allowing the model analysis system to accurately determine the impact of the features of the data points while providing interpretable results.”).
The rationale for combining Nguyen, Redmon, Saralajew, and Ming with Pai is similar to that as applied for claim 7 above.
Claims 2 and 3 are rejected under 35 U.S.C. 103 as being unpatentable over Nguyen in view of Redmon, Saralajew, and Ming, as applied to claims 1, 4, 6, and 8 above, and in further view of Weng et al. (US 7711663 B2), hereafter referred to as Weng.
Regarding claim 2 Nguyen, Redmon, Saralajew, and Ming teaches the limitations of claim 1. Weng, in the same field of neural network classification, teaches the following limitations which the above prior art fails to teach:
wherein the performing performs the prediction and prediction support operation that provides a value of prediction and an interpretation for the value of prediction responsive to the input segments and the similarity vector (Weng, col. 9, lines 40-41, “Then, the amnesic mean is an unbiased estimator of Ex:
PNG
media_image8.png
304
545
media_image8.png
Greyscale
Col. 9, lines 28-29, “Since all the multiplicative factors above are non-negative, we have
W
t
n
≥
0
”, It is well known in the art prior to the filing date of the invention that estimating and predicting can be used interchangeably such that an estimator is interpreted as being used for prediction, which reinforced by Weng, col.5, lines 6-17, “This network incorporates unsupervised learning, supervised learning and reinforcement learning through the developmental process. When the desired action is available to be imposed on the motor layer, the network performs supervised learning. When the desired action is not available but the predicted values of a few top choices of outputs from the network are available, the network is used to perform reinforcement learning by producing the action that has the best predicted value. When none of the desired action nor the predicted values for the top choices of outputs are available, the network performs unsupervised learning and produces an output vector at the output layer based on its recurrent computation”,
In this case, a prediction (also prediction support operation) and interpretation for value of prediction is given by an unbiased estimator using amnesic mean, therefore it is lacking a weight bias. Amnesic mean uses multiplicative factors that are non-negative, therefore the weights will have to be greater than 0.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combined teachings of Nguyen, Redmon, Saralajew, and Ming with the teachings of Weng it provides an explicit technique for producing an unbiased estimator using an amnesic mean in an incremental/online learning setting, which is a predictable improvement when a system relies on estimated statistics during learning and/or interpretation. Wang explains that “each neuron does not have extra space to store all the training samples… [and] uses… mechanisms to update… incrementally,” (Weng, col 7 , line 47) and uses an “amnesic mean technique… which gradually ‘forgets’ old ‘observations’… while keeping the estimator… efficient,” (Weng, col 8 , line 43) making it suitable for incremental estimation in changing environments. A motivation for the combination would have been to provide an incremental and unbiased estimator for interpretation-related quantities.
Regarding claim 3 Nguyen, Redmon, Saralajew, Ming and Weng teaches the limitations of claim 2. Weng further teaches:
wherein the performing performs, by a fully connected layer (Weng, col. 6, lines 63-65, “Given a limited cortical resource, c cells fully connected to input y”)
the prediction and prediction support operation that provides a value of prediction and an interpretation, wherein the interpretation for the value of prediction is provided using only non-negative weights and lacking a weight bias in the fully connected layer (Weng, col. 9, lines 40-41, “Then, the amnesic mean is an unbiased estimator of Ex:
PNG
media_image8.png
304
545
media_image8.png
Greyscale
Col. 9, lines 28-29, “Since all the multiplicative factors above are non-negative, we have
W
t
n
≥
0
”, It is well known in the art prior to the filing date of the invention that estimating and predicting can be used interchangeably such that an estimator is interpreted as being used for prediction, which reinforced by Weng, col.5, lines 6-17, “This network incorporates unsupervised learning, supervised learning and reinforcement learning through the developmental process. When the desired action is available to be imposed on the motor layer, the network performs supervised learning. When the desired action is not available but the predicted values of a few top choices of outputs from the network are available, the network is used to perform reinforcement learning by producing the action that has the best predicted value. When none of the desired action nor the predicted values for the top choices of outputs are available, the network performs unsupervised learning and produces an output vector at the output layer based on its recurrent computation”,
In this case, a prediction (also prediction support operation) and interpretation for value of prediction is given by an unbiased estimator using amnesic mean, therefore it is lacking a weight bias. Amnesic mean uses multiplicative factors that are non-negative, therefore the weights will have to be greater than 0.).
The rationale for combining Nguyen, Redmon, Saralajew, and Ming is similar to that as applied for claim 2 above.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Nguyen in view of Redmon, Saralajew, and Ming, as applied to claims 1, 4, 6, and 8 above, and in further view of Xu et al. (US 20210181931 A1), hereinafter referred to as Xu.
Regarding claim 5, Nguyen, Redmon, Saralajew, and Ming teaches the limitations of claim 1. Xu, in the same field of sequence modeling with prototypes, teaches the following limitation which Nguyen, Redmon, Saralajew and Ming fail to teach:
wherein each of the class prototypes collectively form a class prototype vector that is a latent representation of a prototypical segment (Xu, paragraph [0053], “The prototype layer p compares the latent vector representation e obtained through the encoder network with k prototype vectors”, class prototypes collectively form a class prototype vector to obtain a latent representation. Paragraph [0054] discloses that these prototype vectors are learned through backpropagation of the neural network.).
It would have been obvious to a person of ordinary skill in the art to have incorporated the teachings disclosed by Nguyen, Redmon, Saralajew, and Ming with the teachings disclosed by Xu (i.e., forming a class prototype vector). A motivation for the combination is to provide a way to obtain similarity scores to indicate input sequences with similar prototype embedding (identifying a classification). (Xu, paragraph [0053], “Through appropriate transformations, a vector of similarity scores may be obtained as a=p(e),a.sub.i∈ [0,1], where a.sub.i is the similarity score between the input sequence and the prototype p.sub.i and a.sub.i=1 indicates that the input sequence has identical embedding with prototype p.sub.i.”)
Redmon further teaches:
learned through gradient descent (Redmon, paragraph 39, “The convolutional neural network 110 may be trained in part using training images with associated localization labels and classification labels and be trained in part using training images with associated classification labels that lack localization labels.”, neural network 110 is trained on training images with their associated localization labels and classification labels (class prototypes).
Paragraph 33, “For example, this modified convolutional neural network may be trained with the ImageNet 1000-class classification dataset for 160 epochs using stochastic gradient descent with a starting learning rate of 0:1”, the training is performed with stochastic gradient descent)
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Nguyen in view of Redmon, Saralajew, Ming and Weng, as applied to claims 2 and 3 above, and in further view of Xu.
Regarding claim 13, Nguyen, Redmon, Saralajew, Ming and Weng teaches the limitations of claim 3. Xu teaches the following limitation which Nguyen, Redmon, Saralajew, Ming, and Weng fail to teach:
applying a softmax operation to an output of the fully connected layer to obtain the value of prediction and the interpretation for the value of prediction (Xu, paragraph [0053], “The fully connected layer f with softmax output computes the eventual classification results using the similarity score vector a. The entries in the weight matrix in f may be constrained to be non-negative for better interpretability.”).
It would have been obvious to a person of ordinary skill in the art to have incorporated the teachings disclosed by Nguyen, Redmon, Saralajew, Ming and Weng with the teachings disclosed by Xu (i.e., forming a class prototype vector). A motivation for the combination is to provide a way to obtain similarity scores to indicate input sequences with similar prototype embedding (identifying a classification). (Xu, paragraph [0053], “Through appropriate transformations, a vector of similarity scores may be obtained as a=p(e),a.sub.i∈ [0,1], where a.sub.i is the similarity score between the input sequence and the prototype p.sub.i and a.sub.i=1 indicates that the input sequence has identical embedding with prototype p.sub.i.”)
Claims 9 and 19 is rejected under 35 U.S.C. 103 as being unpatentable over Nguyen in view of Redmon, Saralajew, and Ming, as applied to claims 1, 4, 6, and 8 above, and in further view of Guo et al. (US 20200151250 A1), hereafter referred to as Guo.
Regarding claim 9, Nguyen, Redmon, Saralajew, and Ming teaches the limitations of claim 1. Guo, in the same field of neural network sequence modeling, teaches the following limitation which Nguyen, Redmon, Saralajew, and Ming fail to teach:
wherein the one or more filters comprise multiple filters having different size lengths configured to selectively address different resolutions of the plurality of input segments corresponding to different levels of granularity (Guo, paragraph [0013], “Embodiments of the present invention provide sequence modeling for natural language processing applications using a convolution of kernels of multiple sizes to capture sentence structure at different levels of granularities. This generates a set of feature maps for each position in the sentence that are added together with multi-resolution attention weights to produce the input of a recurrent neural network that generates a context vector.”, in this case filters (kernels) of multiple sizes are used to capture different granularities of input segments (sentence structure). This produces weights that are specifically attuned to multiple resolutions depending on the input.).
It would have been obvious to a person of ordinary skill in the art to have incorporated the teachings disclosed by Nguyen, Redmon, Saralajew, and Ming with the teachings disclosed by Guo (i.e., having multi-resolution attention weights to corresponding levels of granularity). A motivation for the combination is to enable identifying features from input with varying levels of resolution. (Guo, paragraph [0047], “This improves natural language processing tasks such as, e.g., sentiment classification, machine translation, and language modeling. These represent substantive technical fields, and improvements to their ability to consider contextual information at multiple resolutions provide substantial benefits across a wide variety of disciplines.”)
Regarding claim 19, Nguyen, Redmon, Saralajew, and Ming teaches the limitations of claim 1. Guo, in the same field of neural network sequence modeling, teaches the following limitation which Nguyen, Redmon, Saralajew, and Ming fail to teach:
The computer-implemented method of claim 1, wherein the convolutional layer has a plurality of filters of different sizes. (Guo, paragraph [0013], “Embodiments of the present invention provide sequence modeling for natural language processing applications using a convolution of kernels of multiple sizes to capture sentence structure at different levels of granularities. This generates a set of feature maps for each position in the sentence that are added together with multi-resolution attention weights to produce the input of a recurrent neural network that generates a context vector.”, in this case filters (kernels) of multiple sizes are used to capture different granularities of input segments (sentence structure). This produces weights that are specifically attuned to multiple resolutions depending on the input.).
Claim 11 are rejected under 35 U.S.C. 103 as being unpatentable over Nguyen in view of Redmon, Saralajew, Ming, and Weng, as applied to claims 2 and 3 above, and in further view of Bocklet et al. (US 10650807 B2), hereinafter referred to as Bocklet.
Regarding claim 11, Nguyen, Redmon, Saralajew, Ming, and Weng teaches the limitations of claim 2. Bocklet, in the same field of neural network sequence modeling, teaches the following limitation which the above prior art fails to teach:
performing a max pooling operation on the similarity scores to obtain a similarity vector (Bocklet, col. 4, lines 53, “The propagation from state to state along the multiple element state score vector is accomplished by performing a maximum pooling or other down-sampling operation between adjacent scores along the vector of intermediate scores to establish a current multiple element state score vector.”).
It would have been obvious to a person of ordinary skill in the art to have incorporated the teachings disclosed by Nguyen, Redmon, Saralajew, Ming, and Weng with the teachings disclosed by Bocklet (i.e., max pooling similarity scores to create a vector). A motivation for the combination is to create bias values for an input to create the backward operation of a recurrence. (Bocklet, paragraph [0031], “The scores forming the previous multiple element state score vector may be treated as bias values for input through a bias pathway of a neural network accelerator to create the backward operation of the recurrence.”)
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Nguyen in view of Redmon, Saralajew, and Ming, as applied to claims 1, 4, 6, and 8 above, and in further view of Ren et al. (US 20200364504 A1), hereinafter referred to as Ren.
Regarding claim 16, Nguyen, Redmon, Saralajew, and Ming teaches the limitations of claim 1. Ren, in the same field of neural network sequence modeling, teaches the following limitation which the above prior art fails to teach:
wherein said calculating step comprises using a diversity regularization that penalizes small distances between the respective resolution-controllable class prototypes below a diversity regularization threshold distance (Ren, paragraph [0046], “The diversity regularization term may be expressed as: where d.sub.min is a threshold that classifies whether two prototypes are close or not. In some examples, the value of d.sub.min may be set to 1.0 or 2.0. R.sub.d is a soft regularization that exerts a larger penalty on smaller pairwise distances.”).
It would have been obvious to a person of ordinary skill in the art to have incorporated the teachings disclosed by Nguyen, Redmon, Saralajew, and Ming with the teachings disclosed by Ren (i.e., adding a diversity regularizer). A motivation for the combination is to prevent having multiple similar prototypes. (Ren, paragraph [0046], “Having multiple similar prototypes in the explanations can result in confusion and inefficiency in utilizing model parameters. To prevent this, a diversity regularization term may be incorporated that penalizes prototypes that are close to each other.”)
Claims 17 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Nguyen in view of Redmon, Saralajew, and Ming, as applied to claims 1, 4, 6, and 8 above, and in further view of Kuang et al. (Kuang, L., & Zulkernine, M. (2007). Dnids: A dependable network intrusion detection system using the csi-knn algorithm.), hereinafter referred to as Kuang.
Regarding claim 17, Nguyen, Redmon, Saralajew, and Ming teaches the limitations of claim 1. Kuang, in the same field of neural network sequence modeling, teaches the following limitation which the above prior art fails to teach:
The computer-implemented method of claim 1, further comprising performing a control action responsive to the prediction and prediction support operation. (Kuang, page 82, paragraph 1, “The Manager first parses the detector alert and gets the name of the failed detector. Based on the detector name, the Manager locates the affected classifiers from the services table. Then, the Manager sets these classifiers as “unavailable” and switches the maintenance status to “under recovery”.”, Kuang outlines a predictive system that manages a network’s security through machine learning agents or managers. In this case, the system predicts a failure via a detector alert, and performs a control action of identifying and fixing the failed component by deploying Maintenance Agents such as the Manager.
Page 87, paragraph 1, “After the updates, the Manager generates an MA to restore the failed “smtp” classifier. The first available detector other than the detector where the failed classifier was located is selected. Thus, the MA builds the new smtp classifier on detector “\\host1:5555”. After the classifier is established and activated, a recovery-complete message is sent to the Manager and the two tables are updated again.”, the failed component of the network is reestablished).
It would have been obvious to a person of ordinary skill in the art to have incorporated the teachings disclosed by Nguyen, Redmon, Saralajew, and Ming with the teachings disclosed by Kuang and incorporate failed component switching and network security options based off the machine learning predictions made from the previous teachings. A motivation for the combination is to automatically identify failing parts of a network and to replace it, creating improved network security. (Kuang, page 5, paragraph 3, “The intrusion-tolerant mechanism employs mobile agents to check the status of the monitored components and execute recovery tasks. Mobile agents are software programs that move to different hosts and perform tasks. A manager process works with the mobile agents to achieve intrusion tolerance. The process generates two different mobile agents and dispatches them to execute tasks. One mobile agent checks the status of the classifiers and hosts using encoded rules. When it detects a failed component, the agent raises a security alert and reports to the manager process”, the intrusion tolerant system automatically deploys machine learning agents to identify and replace faulty components.)
Regarding claim 18, Nguyen, Redmon, Saralajew, Ming and Kuang teaches the limitations of claim 17. Kuang further teaches:
The computer-implemented method of claim 17, wherein the control action is selected from the group consisting of turning off an impending failing element, swapping out a failed component for another operating component, (Kuang, page 82, paragraph 1, “The Manager first parses the detector alert and gets the name of the failed detector. Based on the detector name, the Manager locates the affected classifiers from the services table. Then, the Manager sets these classifiers as “unavailable” and switches the maintenance status to “under recovery”.”, Kuang outlines a predictive system that manages a network’s security through machine learning agents or managers. In this case, the system predicts a failure via a detector alert, and performs a control action of identifying and fixing the failed component by deploying Maintenance Agents such as the Manager.
Page 87, paragraph 1, “After the updates, the Manager generates an MA to restore the failed “smtp” classifier. The first available detector other than the detector where the failed classifier was located is selected. Thus, the MA builds the new smtp classifier on detector “\\host1:5555”. After the classifier is established and activated, a recovery-complete message is sent to the Manager and the two tables are updated again.”, the failed component of the network is reestablished)
and switching to a secure network. (Kuang, page 76, section 4.2.4, paragraph 2, “The Manager receives various alerts and responds accordingly. For intrusion alerts, the Manager forwards the alerts to the IDS console where a security administrator who is responsible for network security analyzes the alerts and takes corresponding actions. For security alerts, the Manager generates MAs to execute recovery operations. The MA takes the information about the failed classifier from the Manager and restores the classifier in a different operational host, sending the Manager a message as soon as the recovery is finished. The Manager responds to recovery-complete messages by updating the NIDS resource information and executing a reconfiguration.”, when a failure occurs, the system identifies and restores the affected component, afterwards the Network Intrusion Detection System (NIDS) is reconfigured to a more secure state (a state without a failing component).).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Agrawal, M., & Sawhney, K. (2016). Exploring convolutional neural networks for automatic image colorization. Stanford University.
Posada Aguilar, J. D. (2018). Semantics Enhanced Deep Learning Medical Text Classifier (Doctoral dissertation, University of Pittsburgh).
Bulat, A., Tzimiropoulos, G., Kossaifi, J., & Pantic, M. (2019). Improved training of binary networks for human pose estimation and image recognition. arXiv preprint arXiv:1904.05868.
Li, O., Liu, H., Chen, C., & Rudin, C. (2018, April). Deep learning for case-based reasoning through prototypes: A neural network that explains its predictions. In Proceedings of the AAAI conference on artificial intelligence (Vol. 32, No. 1).
Choi, E., Bahadori, M. T., Sun, J., Kulas, J., Schuetz, A., & Stewart, W. (2016). Retain: An interpretable predictive model for healthcare using reverse time attention mechanism. Advances in neural information processing systems, 29.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HYUNGJUN B YI whose telephone number is (703)756-4799. The examiner can normally be reached M-F 9-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed application or proceeding is assigned is (571) 272-4046.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/H.B.Y./Examiner, Art Unit 2124
/USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146