DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 1, 12, and 13 recite the limitations:
configuring a Bayesian meta-model to cooperate with the pretrained machine learning model, the Bayesian meta-model being configured to quantify different kinds of uncertainties associated with the pretrained machine learning model, wherein the Bayesian meta-model comprises a plurality of linear layers attached to different intermediate features of the pretrained machine learning model with a final linear layer generating a Dirichlet distribution;
receiving multiple intermediate features extracted from the pretrained machine learning model as inputs; and
generating a Dirichlet distribution over a probability simplex as output, wherein the Dirichlet distribution is parameterized by the Bayesian meta-model and allows quantification of uncertainty of model prediction.
The bolded terms above have antecedent basis issues. It is also unclear whether the configuring step encapsulates the receiving and generating steps or whether these steps are repeated twice. Appropriate correction is required.
For purposes of examination, terms that are similar will be interpreted as referring to the same element, and the steps will be interpreted as not being repeated (e.g., the intermediate features are only extracted and provided as input once, the Dirichlet distribution is only generated and output once, and the uncertainty quantification is only performed once).
Dependent claims 2-11 and 14-20 fail to cure the deficiencies of the claims from which they depend and are rejected for the same reasons.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 6-9, 12-14, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (US Pat. 11809976) in view of Rhodes (US Pub. 20230334296).
Referring to claim 1, Li discloses A computer-implemented method comprising:
obtaining a pre-trained machine learning model [figs. 2 and 3; col. 7, lines 7-30, line 25; one or more classification models are configured to classify an object, with each classification model trained to provide different outputs from each other depending on the object];
configuring a Bayesian meta-model to cooperate with the pre-trained machine learning model [figs. 2 and 3; col. 7, line 47 – col. 8, line 25; col. 13, line 59 – col. 14, line 3; the classification models generate classifications, uncertainty metrics, and/or object features that are provided to a meta-model, which is a Bayesian Linear Regression (BLR) model], the Bayesian meta-model being configured to quantify different kinds of uncertainties associated with the pre-trained machine learning model [col. 13, line 59 – col. 14, line 3; the meta-model is trained to generate a final classification and a confidence level based on the classifications, the uncertainty metrics, and/or the object features], wherein the Bayesian meta-model comprises a plurality of linear layers attached to different intermediate features of the … machine learning model with a final linear layer generating a … distribution [figs. 2 and 3; col. 6, lines 8-23; col. 7, line 47 – col. 8, line 25; col. 14, lines 4-54; note the meta-model is a BLR model, which means linear regression is used to estimate distributions over parameters and predictions of the classification models; also note the meta-model receives, as input, the classifications, the uncertainty metrics, and/or the object features output by the classification models, where each classification model may be in a hidden (i.e., intermediate) layer];
receiving multiple intermediate features extracted from the pre-trained machine learning model as inputs [figs. 2 and 3; col. 6, lines 8-23; col. 7, line 47 – col. 8, line 25; col. 14, lines 4-54; note the meta-model receives, as input, the classifications, the uncertainty metrics, and/or the object features output by the classification models, where each classification model may be in a hidden (i.e., intermediate) layer];
generating a … distribution over a probability simplex as output, wherein the … distribution is parameterized by the Bayesian meta-model [figs. 2 and 3; col. 6, lines 8-23; col. 7, line 47 – col. 8, line 25; col. 14, lines 4-54; note the meta-model is a BLR model, which means linear regression is used to estimate distributions (e.g., posterior probability distributions) over parameters and predictions of the classification models] and allows quantification of uncertainty of model prediction [col. 13, line 59 – col. 14, line 3; note the confidence level]; and
using the Bayesian meta-model and the pre-trained machine learning model in a downstream task [col. 16, lines 46-55; when the classification models and the meta-model (i.e., together referred to as an ML model) are trained, the trained ML model is used to classify the object].
Li does not appear to explicitly disclose that the distribution is a Dirichlet distribution.
However, Rhodes discloses that the distribution is a Dirichlet distribution [figs. 1 and 2; pars. 29, 30, and 32; multi-view data input into a multi-view data analyzer is processed by an uncertainty estimator to obtain an uncertainty estimate and a final prediction output; in the classification setting, the family of distributions commonly used for this purpose is the Dirichlet distribution, where the support of the Dirichlet distribution in k-dimensions is a k-simplex].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the meta-model taught by Li so that the meta-model generates a Dirichlet distribution as taught by Rhodes, with a reasonable expectation of success. The motivation for doing so would have been because the Dirichlet distribution possesses many useful mathematical properties and can encapsulate cases reflecting confident predictions, conflicting predictions, and out-of-distribution (OOD) predictions [Rhodes, pars. 30 and 32].
Referring to claim 2, Li discloses The computer-implemented method of claim 1, wherein the configuring of the Bayesian meta-model to cooperate with the pretrained machine learning model is performed without modifying the pretrained machine learning model [figs. 2 and 3; col. 6, lines 8-23; col. 7, line 47 – col. 8, line 25; col. 14, lines 4-54; note that the training/configuration of the meta-model does not involve modifying the classification models].
Referring to claim 6, Li discloses The computer-implemented method of claim 1, wherein the final linear layer combines all the intermediate features [figs. 2 and 3; col. 6, lines 8-23; col. 7, line 47 – col. 8, line 25; col. 14, lines 4-54; note the final classification and the confidence level based on the classifications, the uncertainty metrics, and/or the object features (from all of the classification models)].
Referring to claim 7, Rhodes discloses The computer system of claim 1, wherein the Dirichlet distribution comprises a concentrated Dirichlet distribution over the probability simplex corresponding to confident prediction and comprises a diffused Dirichlet distribution corresponding to uncertain predictions [fig. 8A; pars. 28, 33, 34, and 38; in Dirichlet distributions, total vacuity provides an effective means to identify high degrees of epistemic uncertainty; in other words, the total vacuity is directly correlated to uncertainty such that low vacuity (i.e., a concentrated Dirichlet distribution) is associated with a low degree of uncertainty and high vacuity (i.e., a diffused Dirichlet distribution) is associated with a high degree of uncertainty].
Referring to claim 8, Li discloses The computer-implemented method of claim 1, wherein the linear layers consist only of fully connected layers and activation functions [figs. 2 and 3; col. 7, line 47 – col. 8, line 25; col. 13, line 59 – col. 14, line 3; note that the BLR model is, by definition, fully connected because every input connects directly to the output variable through a set of weight coefficients].
Referring to claim 9, Rhodes discloses The computer-implemented method of claim 1, wherein the generating the Dirichlet distribution is based on a loss function and wherein the loss function uses a likelihood term to encourage sharpening of a categorical distribution around a true class on the simplex and uses a KL-divergence term as a regularizer to prevent overconfident prediction, and wherein a hyper-parameter is supplied to balance a trade-off between the sharpening of the categorical distribution and the prevention of the overconfident prediction [pars. 32-34; the multi-view data analyzer creates a categorical distortion form predicted concentration parameters of the Dirichlet distribution, with training to produce high evidence for the ground-truth label class and low evidence for other class assignments using means squared error (MSE) loss to provide a form of baseline regularization by concurrently enforcing variance minimization of the implied Dirichlet distribution; dissonance regularization is used to apply an additional learning constraint via a loss function; uniformed priors are used to enrich evidential distributions by penalizing the generation of the evidence for misclassified data (e.g., to encourage the attribution of low evidence when the model has low prediction confidence)].
Referring to claim 12, Li discloses A computer program product, comprising: one or more tangible computer-readable storage media and program instructions stored on at least one of the one or more tangible computer-readable storage media, the program instructions executable by a processor, the program instructions comprising the claimed steps [fig. 12, processor circuitry 1212, memory 1213, instructions 1232].
Referring to claim 13, Li discloses A system comprising: a memory; and at least one processor, coupled to said memory, and operative to perform operations comprising the claimed steps [fig. 12, processor circuitry 1212, memory 1213, instructions 1232].
Referring to claim 14, see the rejection for claim 2.
Referring to claim 18, see the rejection for claim 6.
Referring to claim 19, see the rejection for claim 7.
Referring to claim 20, see the rejection for claim 8.
Claims 10 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Li and Rhodes in view of Louizos et al. (US Pub. 20200394506).
Referring to claim 10, Li and Rhodes do not appear to explicitly disclose The method of claim 1, further comprising controlling an autonomous vehicle using the Bayesian meta-model in conjunction with the pretrained machine learning model.
However, Louizos discloses The method of claim 1, further comprising controlling an autonomous vehicle using the Bayesian meta-model in conjunction with the pretrained machine learning model [pars. 36, 38, 42, 51, and 52; a machine learning system is used to make classification predictions for autonomous device control (e.g., for autonomous cars); the machine learning system performs probabilistic (Bayesian) modeling and outputs a probability distribution; a control signal is computed (i.e., decision is made) based on a classification result and a confidence value; for example, if the confidence value is low, control (e.g., of the autonomous car) is switched to a more conservative mode such as engaging human control].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify ML model taught by the combination of Li and Rhodes so that the ML model and the confidence level output by the ML model is used to make control decision for an autonomous car as taught by Louizos, with a reasonable expectation of success. The motivation for doing so would have been to improve the accuracy of uncertainty estimates about predictions in safety critical applications such as self-driving cars [Louizos, par. 3].
Referring to claim 11, Li discloses The method of claim 10, further comprising alerting a driver to assume control of the autonomous vehicle in response to a confidence level generated by the Bayesian meta-model in conjunction with the pretrained machine learning model being less than a given threshold [pars. 36, 38, 42, 51, and 52; note the engaging of human control if the confidence is low (relative to an implied threshold)].
Conclusion
The following prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Tsiligkaridis (US Pub. 20210103814) discloses detecting uncertainty using a Bayesian Dirichlet distribution.
Vincon et al. (US Pub. 20250037020) discloses quantifying predictive uncertainty using a Bayesian method and a Dirichlet distribution of K-simplex.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GRACE PARK whose telephone number is (571)270-7727. The examiner can normally be reached M-F 8AM-5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, TAMARA KYLE can be reached at (571)272-4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Grace Park/Primary Examiner, Art Unit 2144