DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-15 are pending.
Claim Objections
Claims 2-4 and 14-15 are objected to because of the following informalities.
In claims 2 and 15, the claim should recite “by performing the following steps:” prior to reciting the specific steps (NB colon after “steps”). This ensures consistency across the claims.
In claim 14, applicant recites “An uncertainty modelling apparatus, comprising: at least one processor, configured to perform the steps: obtaining trained regression model,” which should read: “An uncertainty modelling apparatus, comprising: at least one processor, configured to perform the following steps: obtaining a trained regression model.” This addresses any unintentionally omitted verbiage.
Any dependent claims are objected to by virtue of their dependency. Appropriate correction is required.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 1-15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
In claims 1 and 12, applicant recites “a trained regression model for measured sensor data or image data established to control, to monitor or to analyse a machine, traffic or images in healthcare systems.” There are multiple issues here, given the awkward construction of the claim. First, to what is “established to control” referring? Is it the image that’s controlling or the regression model? Second, are both traffic and images relating to healthcare systems? The specification doesn’t provide much clarity, as seen on page 2, which describes traffic control that could potentially include a healthcare setting. The ambiguity raises the issue of indefiniteness.
In claims 7 and 9, applicant refers to “the subset,” which lacks antecedent basis. Neither the claim nor its parent refers to any “subset.”
In claim 11, applicant refers to “the digital computer,” which lacks antecedent basis. Neither the claim nor its parent refers to any “digital computer”
In claim 14, applicant recites “wherein the trained regression model is established to control, to monitor or to analyse a machine, traffic or images in healthcare systems.” Are both traffic and images relating to healthcare systems? The specification doesn’t provide much clarity, as seen on page 2, which describes traffic control that could potentially include a healthcare setting. The ambiguity raises the issue of indefiniteness.
Any dependent claims are rejected by virtue of their dependency. Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception without significantly more.
Step 1 (The Statutory Categories): Is the claim to a process, machine, manufacture, or composition of matter? MPEP 2106.03.
Per Step 1, claim 1 is to a method (i.e., a process), claim 11 to a computer program product (i.e., an article), claim 12 to an apparatus (i.e., a machine), and claim 14 also to an apparatus (i.e., a machine). Thus, the claims are directed to statutory categories of invention. However, the claims are rejected under 35 U.S.C. 101 because they are directed to an abstract idea, a judicial exception, without reciting additional elements that integrate the judicial exception into a practical application.
The analysis proceeds to Step 2A Prong One.
(Examiner notes that claim 11, while dependent on claim 1, is treated as an independent claim in the analysis below, since it’s a separate statutory category. Regardless of how it’s analyzed, whether independent or dependent, the claim’s still ineligible.)
Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon? MPEP 2106.04.
The abstract idea of claims 1, 11, and 12 and is (claim 1 being representative, claim 11 inheriting the abstract idea from claim 1):
obtaining the trained regression model, training data which were applied to train the regression model, and an empirical variance determined by the regression model applying the training data as input data;
generating an uncertainty layer in the trained regression model based on the training data, and the empirical variance;
obtaining the measured sensor data or image data as input data;
outputting the prediction by processing the input data in the trained regression model; and
outputting an uncertainty value of the prediction by processing the input data by a feature extractor model and subsequently by the uncertainty layer, wherein the feature extractor model comprises all but the last layers of the regression model.
The abstract idea of claim 14 is:
obtaining trained regression model, training data which were applied to train the trained regression model, and an empirical variance determined by the regression model applying the training data as input data;
generating an uncertainty layer in the trained regression model based on the training data, and the empirical variance; and
outputting an enhanced trained regression model which comprises the trained regression model and the uncertainty layer,
wherein the trained regression model is established to control, to monitor or to analyse a machine, traffic or images in healthcare systems.
The abstract idea steps italicized above are those which could be performed mentally, including with pen and paper. At a high level, the steps describe receiving information (obtaining [a]/the trained regression model), analyzing it (generating an uncertainty layer), and outputting results (outputting an uncertainty value or outputting an enhanced trained regression model), in a manner akin to Electric Power Group. The claims do not recite the specifics of the underlying technological architecture or solution; instead, they describe abstract steps associated with: obtaining information, generating an output, and evaluating the output (claims 1, 11, and 12); or obtaining information and generating an output (claim 14). the If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, including observations, evaluations, judgements, and/or opinions, then it falls within the Mental Processes – Concepts Performed in the Human Mind grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Additionally and alternatively, the abstract idea steps italicized above, which include – claims 1, 11, and 12: obtaining the trained regression model; generating an uncertainty layer in the trained regression model; processing the input data in the trained regression model; and outputting an uncertainty value of the prediction by processing the input data by a feature extractor model and subsequently by the uncertainty layer, wherein the feature extractor model comprises all but the last layers of the regression model; claim 14: obtaining [a] trained regression model; generating an uncertainty layer in the trained regression model; outputting an enhanced trained regression model which comprises the trained regression model and the uncertainty layer – describe mathematical calculations pertaining to evaluating an uncertainty value of a prediction, which constitutes a process that, under its broadest reasonable interpretation, covers mathematical concepts. If a claim limitation, under its broadest reasonable interpretation, covers mathematical concepts, including mathematical relationships, mathematical formulas or equations, mathematical calculations, then it falls within the Mathematical Concepts grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application? MPEP 2106.04.
This judicial exception is not integrated into a practical application because the additional elements are merely instructions to apply the abstract idea to a computer, as described in MPEP 2106.05(f).
Claim 1 recites the following additional elements: computer-implemented.
Claim 11 recites the following additional elements: computer program product comprising a computer readable hardware storage device having computer readable program code stored therein, the program code executable by a processor of a computer system […] when the product is run on the digital computer.
Claim 12 recites the following additional elements: at least one processor.
Claim 14 recites the following additional elements: at least one processor.
These elements are merely instructions to apply the abstract idea to a computer, per MPEP 2106.05(f). Applicant has only described generic computing elements in their specification, as seen in on pages 9-10 of applicant’s specification as filed.
Further, the combination of these elements is nothing more than a generic computing system applied to the tasks of the abstract idea. Because the additional elements are merely instructions to apply the abstract idea to a generic computing system, they do not integrate the abstract idea into a practical application, when viewed in combination. See MPEP 2106.05(f).
Therefore, per Step 2A Prong Two, the additional elements, alone and in combination, do not integrate the judicial exception into a practical application. The claim is directed to an abstract idea.
Step 2B (The Inventive Concept): Does the claim recite additional elements that amount to significantly more than the judicial exception? MPEP 2106.05.
Step 2B involves evaluating the additional elements to determine whether they amount to significantly more than the judicial exception itself.
The examination process involves carrying over identification of the additional element(s) in the claim from Step 2A Prong Two and carrying over conclusions from Step 2A Prong Two pertaining to MPEP 2106.05(f).
The additional elements and their analysis are therefore carried over: applicant has merely recited elements that facilitate the tasks of the abstract idea, as described in MPEP 2106.05(f).
Further, the combination of these elements is nothing more than a generic computing system applied to the tasks of the abstract idea. When the claim elements above are considered, alone and in combination, they do not amount to significantly more.
Therefore, per Step 2B, the additional elements, alone and in combination, are not significantly more. The claims are not patent eligible.
The analysis takes into consideration all dependent claims as well:
Dependent claims 2-10, 13, and 15 recite additional abstract steps and/or information that further narrow the abstract idea(s) above. There are no further additional elements to consider, beyond those noted above. This simple narrowing of the abstract idea doesn’t integrate it into practical application or add significantly more, and the abstract idea groupings highlighted previously still apply.
Accordingly, claims 1-15 are rejected under 35 USC § 101 as being directed to non-statutory subject matter.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 5-8, and 10-13 are rejected under 35 U.S.C. 103 as being unpatentable over Willers (US 20200410364) in view of Boettcher (US 20160180220) and Richmond (US 20190362226).
Claims 1 and 12
Willers discloses:
[Claim 1: A computer-implemented method for automatically quantifying an uncertainty of a prediction provided by a trained regression model for measured sensor data or image data established to control, to monitor or to analyse a machine, traffic or images in healthcare systems {computer-implemented method for automatically quantifying an uncertainty of a prediction provided by a trained regression model described in [0002]: The present invention is directed to a method for estimating a global uncertainty of output data of a computer implemented neural network, a computer program, a computer readable storage device and an apparatus which is arranged to perform the method. trained regression model also described in [0023]: This approach has the advantage, that uncertainties are created by design. Therefore the second measure can easily be extracted. Additionally, the training process can be split in two by separately training the neural network as feature extractor and the Gaussian process as regressor, especially for classification and regression problems. This reduces the complexity with respect to Bayesian neural networks. measured sensor data or image data established to control, to monitor or to analyse a machine, traffic or images in healthcare systems described in [0067]: Possible areas of application are image recognition (camera images, radar images, lidar images, ultrasound images and especially combinations of these), noise classification and much more. They can be used for security applications (home security), applications in the automotive sector, in the space and aviation industry, for shipping and rail traffic.}, the method comprising:]
[Claim 12: An assistance apparatus for automatically quantifying an uncertainty of a prediction provided by a trained regression model for measured sensor data or image data, established to control, to monitor or to analyse a machine, traffic, or images in healthcare systems {See previous citation to [0002], which also describes an apparatus. Also see previous citations to [0023] and [0067].}, comprising: at least one processor {See [0053].}, configured to perform the steps:]
obtaining the trained regression model, training data which were applied to train the regression model, {obtaining the trained regression model, training data which were applied to train the regression model described in [0092]: Module 210 receives input data 201 and determines the first measure 211, showing whether the current input data follows the same distribution as the training data set. Also see previous citation to [0023], which described the trained regression model.};
generating an uncertainty layer in the trained regression model based on the training data, and the empirical variance {generating an uncertainty layer in the trained regression model based on the training data, and the empirical variance described in [0095]: Module 240 determines the global uncertainty 241 based on the first 211, second 221 and third 231 measure. layer(s) also described in [0011]: The main neural network can be a deep neural network, wherein a deep neural network normally comprises at least an input layer, an output layer and at least one hidden layer.};
obtaining the measured sensor data or image data as input data {obtaining the measured sensor data or image data as input data described in [0099]: The different sensors 311, 321, 331 generate sensor data, like images 312, radar signatures 322 and 3D-scans from the lidar sensor 331. This sensor data is fused within block 302. In this example all sensor data 312, 322, 332 is fed into a neural network module 303 which is designed to output the predictions 305 of a neural network and uncertainties 304 about this prediction 305.};
outputting the prediction by processing the input data in the trained regression model {outputting the prediction by processing the input data in the trained regression model described in [0099]: In this example all sensor data 312, 322, 332 is fed into a neural network module 303 which is designed to output the predictions 305 of a neural network and uncertainties 304 about this prediction 305.}; and
outputting an uncertainty value of the prediction by processing the input data by a feature extractor model and subsequently by the uncertainty layer {outputting an uncertainty value of the prediction by processing the input data subsequently by the uncertainty layer described in [0011]: The main neural network can be a deep neural network, wherein a deep neural network normally comprises at least an input layer, an output layer and at least one hidden layer. Also see previous citation to [0099]. by a feature extractor model described in [0024]: The approach of using a neural network as feature extractor, which feeds its output to a Gaussian process, is also known as deep kernel learning in the literature if the neural network and the Gaussian process are trained combined.}.
Willers doesn’t explicitly disclose, however, Boettcher, in a similar field of endeavor directed to adaptively updating equipment models, teaches:
an empirical variance determined by the regression model applying the training data as input data {an empirical variance determined by the regression model applying the training data as input data described in [0004], [0086], [0098]: [0004] In some embodiments, adaptively updating the predictive model includes generating a new set of model coefficients for the predictive model, determining whether the new set of model coefficients improves a fit of the predictive model to a set of operating data relative to a previous set of model coefficients used in the predictive model, and replacing the previous set of model coefficients with the new set of model coefficients in the predictive model in response to a determination that the new set of model coefficients improves the fit of the predictive model. [0086] In some embodiments, model update module 328 refits the model coefficients in response to a determination that the model coefficients have likely changed. Refitting the model coefficients may include using operating data from BAS database 312 to retrain the predictive model and generate a new set of model coefficients {circumflex over (β)}.sub.2. The predictive model may be retained using new operating data (e.g., new variables, new values for existing variables, etc.) gathered in response to a determination that the model coefficients have likely changed or existing data gathered prior to the determination. [0098] Balance point module 418 may be configured to find an optimal balance point for a calculated variable (e.g., a variable based on an enthalpy value calculated in enthalpy module 416, an outdoor air temperature variable, etc.). Balance point module 418 may determine a base value for the variable for which the estimated variance of the regression errors is minimized.}
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to modify Willers to include the features of Boettcher. Given that Willers is directed to utilizing neural networks for classification, one of ordinary skill in the art would have been motivated to look to Boettcher, in order to improve a fit of the predictive model to a set of operating data relative to a previous set of model coefficients used in the predictive model {See [0004] of Boettcher.}.
The combination of Willers and Boettcher doesn’t explicitly teach, however, Richmond, in a similar field of endeavor directed to pre-trained convolutional neural networks, teaches:
wherein the feature extractor model comprises all but the last layers of the regression model {wherein the feature extractor model comprises all but the last layers of the regression model described in [0016]: To design a deep learning architecture, the present methods and systems may implement various transfer learning strategies. Examples of such strategies include, but are not limited to: Treating the convolutional neural network as a fixed feature extractor: Given a convolutional neural network pre-trained on ImageNet, the last fully connected layer may be removed, then the convolutional neural network may be treated as a fixed feature extractor for the new dataset.}.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to modify Willers and Boettcher to include the features of Richmond. Given that Willers is directed to utilizing neural networks for classification, one of ordinary skill in the art would have been motivated to look to Richmond, in order to improve transfer learning or domain adaptation when dealing with a domain with image characteristics different than those found in ImageNet or other large dataset of images {See [0016] of Richmond.}.
Claim 5
Boettcher further teaches: wherein the training data is a subset of training data comprising less training data than the entire training dataset used to train the regression model {wherein the training data is a subset of training data comprising less training data than the entire training dataset used to train the regression model described in [0099]: Regression period module 422 may be configured to determine periods of time that can be reliably used for model regression by model generator module 314 and data synchronization module 404. In some embodiments, regression period module 422 uses sliding data windows to identify a plurality of data samples for use in a regression analysis. For example, regression period module 422 may define regression periods T.sub.1 and T.sub.2 and identify a plurality of data samples within each regression period. In some embodiments, regression period module 422 defines regression period T.sub.2 as a fixed-duration time window (e.g., one month, one week, one day, one year, one hour, etc.) ending at the current time. Regression period module 422 may define regression period T.sub.1 as a fixed-duration time window occurring at least partially prior to regression period T.sub.2 (e.g., ending at the beginning of regression period T.sub.2, ending within regression period T.sub.2, etc.). As time progresses, regression period module 422 may iteratively redefine regression periods T.sub.1 and T.sub.2 based on the current time at which the regression is performed. Sliding data windows are described in greater detail with reference to FIG. 5.}.
The motivation and rationale to include the additional features of Boettcher is the same as set forth previously.
Claim 6
Boettcher further teaches: wherein the subset of training data comprises uniformly distributed random samples of the training data {wherein the subset of training data comprises uniformly distributed random samples of the training data described in [0110]: The K-fold cross validation method is configured to randomly partition the historical data provided to model generator module 314 into K number of subsamples for testing against the baseline model. In other embodiments, a repeated random sub-sampling process (RRSS), a leave-one-out (LOO) process, a combination thereof, or another suitable cross-validation routine may be used by cross-validation module 408.}.
The motivation and rationale to include the additional features of Boettcher is the same as set forth previously.
Claim 7
Boettcher further teaches: wherein the subset of training data comprises data samples representing cluster centers of clusters resulting from a cluster analysis on the entire training dataset {wherein the subset of training data comprises data samples representing cluster centers of clusters resulting from a cluster analysis on the entire training dataset described in [0093]: [0093] According to another exemplary embodiment, outlier analysis module 410 can be configured to conduct a cluster analysis. The cluster analysis may be used to help identify and remove unreliable data points. For example, a cluster analysis may identify or group operating states of equipment (e.g., identifying the group of equipment that is off). A cluster analysis can return clusters and centroid values for the grouped or identified equipment or states. The centroid values can be associated with data that is desirable to keep rather than discard. Cluster analyses can be used to further automate the data clean-up process because little to no configuration is required relative to thresholding.}.
The motivation and rationale to include the additional features of Boettcher is the same as set forth previously.
Claim 8
Willers further discloses: wherein the uncertainty value of the prediction is a variance of the prediction {wherein the uncertainty value of the prediction is a variance of the prediction described in [0071]: The average of the sampled outputs is then used as prediction from the model while the variance of the predictions can be used for computing an uncertainty measure.}.
Claim 10
Willers further discloses: wherein the trained regression model is applied for condition monitoring or quality control or image recognition in a manufacturing process or in autonomous driving or in healthcare support {See previous citation to [0067].}.
Claim 11
Willers further discloses: A computer program product comprising a computer readable hardware storage device having computer readable program code stored therein, the program code executable by a processor of a computer system to implement a method of claim 1 when the product is run on the digital computer {[0055] Additionally, in accordance with the present invention, a computer program is provided with program code which may be stored on a machine-readable carrier or storage medium such as a semiconductor memory, a hard disk memory or an optical memory and which is used to carry out, implement and/or control the steps of the process according to one of the forms of execution described above, in particular when the program product or program is executed on a computer or a device, is also advantageous. Also see [0053].}.
Claim 13
Willers further discloses: wherein the assistance apparatus is installed and/or deployed on the device or system to which sensors are deployed providing the measured sensor data or image data, or on a cloud, or on an edge device {sensors described in [0044]: The input data representing that area, can for example originate from sensors that scan the area. For example a camera, a lidar, a radar or an ultrasonic sensor. Also see previous citation to [0067].}
Claims 2-3 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Willers, Boettcher, and Richmond, further in view of Adams (US 20140358831).
Claim 2
Willers further discloses: wherein the uncertainty layer is generated comprising: a. splitting the regression model into a linear model comprising a last layer of the regression model, and the feature extractor model {See previous citation to [0023].};
The combination of Willers, Boettcher, and Richmond doesn’t explicitly teach, however, Adams, in a similar field of endeavor directed to optimizing machine learning techniques, teaches:
b. determining latent representations by applying the training data to the feature extractor {determining latent representations by applying the training data to the feature extractor described in [0083], which details training inputs pushed through and projections as latent representations}; and c. generating the uncertainty layer by combining the latent representations with the empirical variance {generating the uncertainty layer by combining the latent representations with the empirical variance described in [0083], which also details combining latent representations and predictive variance.}.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to modify the combination of Willers, Boettcher, and Richmond to include the features of Adams. Given that Willers is directed to utilizing neural networks for classification, one of ordinary skill in the art would have been motivated to look to Adams, in order to improve the robustness and performance of conventional Bayesian optimization techniques {See [0066] of Adams.}.
Claim 3
Boettcher further teaches: wherein the latent representations are combined with the empirical variance according to an Ordinary Least Square Model {wherein the latent representations are combined with the empirical variance according to an Ordinary Least Square Model described in [0066]: Model generator module 314 generate an ordinary least squares (OLS) estimate of the model coefficients {circumflex over (β)} by finding the model coefficient vector β that minimizes the RSS function. According to various exemplary embodiments, other methods than RSS and/or OLS may be used (e.g., weighted linear regression, regression through the origin, a principal component regression (PCR), ridge regression (RR), partial least squares regression (PLSR), etc.) to generate the model coefficients.}.
The motivation and rationale to include the additional features of Boettcher is the same as set forth previously.
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Willers, Boettcher, Richmond, and Adams, further in view of Dasgupta (US 20210042619).
Claim 4
The combination of Willers, Boettcher, Richmond, and Adams doesn’t explicitly teach, however, Dasgupta, in a similar field of endeavor directed to deep neural network techniques, teaches: wherein the latent representations are combined with the empirical variance as a Gaussian Process model with a linear kernel defined on the latent representations {wherein the latent representations are combined with the empirical variance as a Gaussian Process model with a linear kernel defined on the latent representations described in [0052]-[0053]: [0052] By contrast, finite rank deep kernel learning expresses a composite kernel as a sum of multiple simpler dot kernels, which cover the space of finite rank Mercer kernels 104, as depicted in FIG. 1. The composite kernels may be expressed as follows: K(x,y)=Σ.sub.i=1.sup.Rϕ.sub.i(x)ϕ.sub.i(y), [0053] where ϕ.sub.i(x)'s form a set of orthogonal embeddings, which are learnt by a deep neural network. By expressing the composite kernel in this fashion, one can show that the possible set of kernels would become richer than the existing kernels adopted in conventional deep kernel learning approaches.}.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to modify the combination of Willers, Boettcher, Richmond, and Adams to include the features of Dasgupta. Given that Willers is directed to utilizing neural networks for classification, one of ordinary skill in the art would have been motivated to look to Dasgupta, in order to facilitate improved performance, such as faster operation, lower processing requirements, and less memory usage, as compared to conventional machine learning methods {See [0028] of Dasgupta.}.
Claims 9 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Willers, Boettcher, and Richmond, further in view of Qiu (US 20200372327).
Claim 9
Willers further discloses: for the measured sensor or image data {See previous citation to [0067].}
The combination of Willers, Boettcher, and Richmond doesn’t explicitly teach, however, Qiu, in a similar field of endeavor directed to quantifying the predictive uncertainty of neural networks, teaches: wherein the uncertainty layer comprises an uncertainty core element, which is calculated depending on the subset of training data during generation of the uncertainty layer and wherein the calculated uncertainty core element is reused during outputting the uncertainty value of the prediction {accomplished via quantification of uncertainty, where evaluation and subsequent application represents an uncertainty core element, as described in [0027], [0029]: [0027] In operation, the RIO process quantifies uncertainty in a NN's predictions. And in the following exemplary point prediction model, RIO calibrates the point predictions of the NN to make them more accurate. Consider a training dataset custom-character=(custom-character,y)={(x.sub.i,y.sub.i)}.sub.i=1.sup.n, and a pretrained standard NN that outputs a point prediction ŷ.sub.i given x.sub.i. RIO addresses the uncertainty problem by modeling the residuals between observed outcomes y and NN predictions ŷ using GP with a composite kernel. The RIO procedure is captured in Algorithm 1 and described in detail below. [0029] In the deployment or real-world application phase, the (NN.sub.PT) 10 is applied S15 to real-world data (x.sub.RW) from database 20 and produces predicted outcomes (x.sub.RW,y.sub.RW) 25. The trained GP (GP.sub.T) 15 is applied S20 to provide reliable uncertainty estimates 30 for the NN.sub.PT predicted outcomes which are then used to calibrate S25 the predictions to provide calibrate predictions (x.sub.RW, Y.sub.RW′) 35. Overall, the RIO method transforms the standard neural network regression model into a reliable probabilistic estimator.} for the measured sensor data or image data.}.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to modify the combination of Willers, Boettcher, and Richmond to include the features of Qiu. Given that Willers is directed to utilizing neural networks for classification, one of ordinary skill in the art would have been motivated to look to Qiu, in order to facilitate quantifying point-prediction uncertainty in standard NNs and other prediction models {See [0007] of Qiu.}.
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Willers in view of Boettcher.
Claim 14
Willers disclose:
An uncertainty modelling apparatus {See previous citation to [0002] above.}, comprising: at least one processor {See previous citation to [0053] above.}, configured to perform the steps:
obtaining trained regression model, training data which were applied to train the trained regression model {obtaining trained regression model, training data which were applied to train the regression model described in [0092]: Module 210 receives input data 201 and determines the first measure 211, showing whether the current input data follows the same distribution as the training data set. trained regression model described in [0023]: This approach has the advantage, that uncertainties are created by design. Therefore the second measure can easily be extracted. Additionally, the training process can be split in two by separately training the neural network as feature extractor and the Gaussian process as regressor, especially for classification and regression problems. This reduces the complexity with respect to Bayesian neural networks.};
generating an uncertainty layer in the trained regression model based on the training data, and the empirical variance {generating an uncertainty layer in the trained regression model based on the training data, and the empirical variance described in [0095]: Module 240 determines the global uncertainty 241 based on the first 211, second 221 and third 231 measure. layer(s) also described in [0011]: The main neural network can be a deep neural network, wherein a deep neural network normally comprises at least an input layer, an output layer and at least one hidden layer.}; and
outputting an enhanced trained regression model which comprises the trained regression model and the uncertainty layer {outputting an enhanced trained regression model which comprises the trained regression model and the uncertainty layer described in [0099]: In this example all sensor data 312, 322, 332 is fed into a neural network module 303 which is designed to output the predictions 305 of a neural network and uncertainties 304 about this prediction 305.},
wherein the trained regression model is established to control, to monitor or to analyse a machine, traffic or images in healthcare systems {wherein the trained regression model is established to control, to monitor or to analyse a machine, traffic or images in healthcare systems described in [0067]: Possible areas of application are image recognition (camera images, radar images, lidar images, ultrasound images and especially combinations of these), noise classification and much more. They can be used for security applications (home security), applications in the automotive sector, in the space and aviation industry, for shipping and rail traffic.}.
Willers doesn’t explicitly disclose, however, Boettcher, in a similar field of endeavor directed to adaptively updating equipment models, teaches:
an empirical variance determined by the regression model applying the training data as input data {an empirical variance determined by the regression model applying the training data as input data described in [0004], [0086], [0098]: [0004] In some embodiments, adaptively updating the predictive model includes generating a new set of model coefficients for the predictive model, determining whether the new set of model coefficients improves a fit of the predictive model to a set of operating data relative to a previous set of model coefficients used in the predictive model, and replacing the previous set of model coefficients with the new set of model coefficients in the predictive model in response to a determination that the new set of model coefficients improves the fit of the predictive model. [0086] In some embodiments, model update module 328 refits the model coefficients in response to a determination that the model coefficients have likely changed. Refitting the model coefficients may include using operating data from BAS database 312 to retrain the predictive model and generate a new set of model coefficients {circumflex over (β)}.sub.2. The predictive model may be retained using new operating data (e.g., new variables, new values for existing variables, etc.) gathered in response to a determination that the model coefficients have likely changed or existing data gathered prior to the determination. [0098] Balance point module 418 may be configured to find an optimal balance point for a calculated variable (e.g., a variable based on an enthalpy value calculated in enthalpy module 416, an outdoor air temperature variable, etc.). Balance point module 418 may determine a base value for the variable for which the estimated variance of the regression errors is minimized.}
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to modify Willers to include the features of Boettcher. Given that Willers is directed to utilizing neural networks for classification, one of ordinary skill in the art would have been motivated to look to Boettcher, in order to improve a fit of the predictive model to a set of operating data relative to a previous set of model coefficients used in the predictive model {See [0004] of Boettcher.}.
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Willers and Boettcher, further in view of Adams.
Claim 15
Willers further discloses: wherein the uncertainty layer is generated by the at least one processor by performing the steps: a. splitting the regression model into a linear model comprising a last layer of the regression model, and the feature extractor model {See previous citation to [0023].};
The combination of Willers and Boettcher doesn’t explicitly teach, however, Adams, in a similar field of endeavor directed to optimizing machine learning techniques, teaches:
b. determining latent representations by applying the training data to the feature extractor {determining latent representations by applying the training data to the feature extractor described in [0083], which details training inputs pushed through and projections as latent representations}; and c. generating the uncertainty layer by combining the latent representations with the empirical variance {generating the uncertainty layer by combining the latent representations with the empirical variance described in [0083], which also details combining latent representations and predictive variance.}.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to modify the combination of Willers and Boettcher to include the features of Adams. Given that Willers is directed to utilizing neural networks for classification, one of ordinary skill in the art would have been motivated to look to Adams, in order to improve the robustness and performance of conventional Bayesian optimization techniques {See [0066] of Adams.}.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
“Neural Network-Based Uncertainty Quantification: A Survey of Methodologies and Applications” (NPL attached), which teaches: Uncertainty quantification plays a critical role in the process of decision making and optimization in many fields of science and engineering. The field has gained an overwhelming attention among researchers in recent years resulting in an arsenal of different methods. Probabilistic forecasting and in particular prediction intervals (PIs) are one of the techniques most widely used in the literature for uncertainty quantification. Researchers have reported studies of uncertainty quantification in critical applications such as medical diagnostics, bioinformatics, renewable energies, and power grids. The purpose of this survey paper is to comprehensively study neural network-based methods for construction of prediction intervals.
US 20190258949, which teaches: A recommendation engine may generate a recommendation in response to user interactions and executed operations in a system. The recommendation may be determined according to a number of factors including, but not limited to, an object affinity and a user affinity. The recommendation may include one or more of a recommendation to use an object and a recommendation for taking one or more actions. The recommendation may be provided to a user if the recommendation satisfies a confidence threshold. Recommendations provided by the recommendation engine are tracked to determine if the user accepted or rejected the recommendations. User history of accepting or rejecting recommendations may be utilized to train the recommendation engine for future recommendations and to build a user profile in a user database.
US 20220277197, which teaches: Methods and systems for language processing include augmenting an original training dataset to produce an augmented dataset that includes a first example that includes a first scrambled replacement for a first word and a definition of the first word, and a second example that includes a second scrambled replacement for the first word and a definition of an alternative to the first word. A neural network classifier is trained using the augmented dataset.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOHN SAMUEL WASAFF whose telephone number is (571)270-5091. The examiner can normally be reached Monday through Friday 8:00 am to 6:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, SARAH MONFELDT can be reached at (571) 270-1833. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
JOHN SAMUEL WASAFF
Primary Examiner
Art Unit 3629
/JOHN S. WASAFF/Primary Examiner, Art Unit 3629