Prosecution Insights
Last updated: October 02, 2026
Application No. 18/822,684

DEVICE AND METHOD FOR DETERMINING AN OUTPUT SIGNAL AND A CONFIDENCE IN THE DETERMINED OUTPUT SIGNAL

Non-Final OA §101§103
Filed
Sep 03, 2024
Priority
Sep 12, 2023 — EU 23 19 6918.9
Examiner
DIEP, DUY T
Art Unit
Tech Center
Assignee
Robert Bosch GmbH
OA Round
1 (Non-Final)
37%
Grant Probability
At Risk
1-2
OA Rounds
2y 3m
Est. Remaining
61%
With Interview

Examiner Intelligence

Grants only 37% of cases
37%
Career Allowance Rate
13 granted / 35 resolved
-22.9% vs TC avg
Strong +24% interview lift
Without
With
+23.7%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
18 currently pending
Career history
65
Total Applications
across all art units

Statute-Specific Performance

§101
29.0%
-11.0% vs TC avg
§103
60.5%
+20.5% vs TC avg
§102
2.8%
-37.2% vs TC avg
§112
7.7%
-32.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 35 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title. Claims 1-4, 6-12 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1, Step 1: Claim 1 recites a method, one of the four statutory categories of patentable subject matter. Step 2A, Prong I: Claim 1 further recites the limitations of: “determining, …, a feature representation of the sensor signal, …” The limitation recites a mental process. A person can mentally determine a feature representation of the sensor signal because the feature representation is just a structured set of numerical values or visual formats. For example, given an image (sensor signal) a person can determine feature representation such as the geometry or color or pixel of the image. “providing a predictive posterior distribution or an argument of a maximum of the predictive posterior distributions as the first element, wherein the predictive posterior distribution is determined based on a posterior distribution of sets of weights of the head and a likelihood of the feature representation given a set of weights of the head” The limitation recites a mental process as well as a mathematical concept. The limitation recites providing a predictive posterior distribution based on a posterior distribution of sets of weights of the head and a likelihood of the feature representation, wherein the posterior distribution of sets of weights constitutes a mathematical concept according to the formula at page 8 and at page 8 line 13 within the Specification “In the formula, the predictive posterior distribution is defined in terms of the input x …”, and the likelihood of the feature representation also constitutes a mathematical concept according to the formula at page 8 and page 8 lines 1-4 “a likelihood of the feature representation given the sampled set of weights can be conducted according to the formula: …”. A person can mentally or manually solve this formula. “sampling a set of weights from the posterior distribution of sets of weights of the head” The limitation recites a mental process as well as a mathematical concept. The process of sampling a set of weights from the posterior distribution constitutes a mathematical concept because it requires a selection of sampled values from a set of values, wherein such selection of values can be mentally performed within a human’s mind. “determining a likelihood ratio for possible classifications or regression results by dividing a value of the predictive posterior distribution at the feature representation by a likelihood of the feature representation given the sampled set of weights” The limitation recites a mental process as well as a mathematical concept. The limitation recites a calculation by performing a division between a distribution and a likelihood value, which constitutes a division between two values and such operation can be mentally or manually performed within a human’s mind. “determining the confidence interval as the possible classes or regression results for which the likelihood ratio is equal to or below a predefined threshold and providing the confidence interval or a value characterizing the width of the confidence interval as the second element” The limitation recites a mental process as well as a mathematical concept. The limitation recites a calculation of the likelihood ratio that corresponds to a mathematical concept as well as a mental process and further compares it with a confidence interval based on a predefined threshold, wherein such comparison is just a mental evaluation between two set of values that can be performed within a human’s mind. Step 2A, Prong II: Claim 1 recites the following additional elements: “… a sensor signal …” These additional elements are a high-level recitation of generic computer components used as a tool and does not provide integration into a practical application. “… wherein the first element characterizes a classification or a regression result of a sensor signal, and the second element characterizes a confidence interval of likely classifications or regression results, wherein the first element and second element are determined by an early-exit neural network (EENN)” This additional element recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f) and does not provide integration into a practical application. The limitation recites the black-box application of using an early-exit neural network (EENN), and using input sensor signal to determine a result and evaluate the result, which are just black-box application of machine learning model to obtain a result and does not provide improvement toward a machine learning algorithm or any computer elements. “determining, by the early-exit neural network, a feature representation of the sensor signal, …” This additional element recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f) and does not provide integration into a practical application. The limitation recites the black-box application of using the early-exit neural network to determine a feature representation, without providing specific details of an unconventional neural network algorithm or neural network architecture to obtain the feature representation as an improvement process. “… a feature representation of the sensor signal, which is provided to a head of the early-exit neural network” This additional element recites an additional element of an insignificant extra-solution activity as identified in MPEP 2106.05(g) of mere data gathering, and does not provide integration into a practical application. The limitation simply recites how input data is obtained for a neural network to process it. The data is obtained at the head of the neural network, which is just a software layer of the neural network for calculating the result. Step 2B: When considered individually or in combination, the additional limitations and elements of claim 1 do not amount to significantly more than the judicial exception for the same reasons discussed above as to why the additional limitations do not integrate the abstract idea into a practical application. The additional elements outlined in Step 2A performing functions as designed simply accomplish execution of the abstract ideas. The additional element “… a sensor signal …” is a high-level recitation of generic computer components used as a tool and does not amount to significantly more than the judicial exception for the same reasons discussed above as to why the additional limitations do not integrate the abstract idea into a practical application. The additional element “… wherein the first element characterizes a classification or a regression result of a sensor signal, and the second element characterizes a confidence interval of likely classifications or regression results, wherein the first element and second element are determined by an early-exit neural network (EENN)” recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not amount to significantly more than the judicial exception for the same reasons discussed above. The additional element “determining, by the early-exit neural network, a feature representation of the sensor signal, …” recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not amount to significantly more than the judicial exception for the same reasons discussed above. The additional element “… a feature representation of the sensor signal, which is provided to a head of the early-exit neural network” recites an additional element of a well-understood, routine, conventional activity as identified in MPEP 2106.05(d) of receiving or transmitting data over a network and does not provide integration into a practical application or amount to significantly more than the judicial exception. In conclusions from above for the elements considered as a mental process, elements reciting high-level recitation of generic computer components used as a tool, and elements reciting a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f) are carried over and do not provide significantly more than the abstract idea. Looking at the limitations in combination and the claims as a whole does not change this conclusion and the claim is ineligible. Therefore, additional limitations of claim 1 do not amount to significantly more than the judicial exception. Thus, claim 1 recites abstract ideas with additional elements rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 1 is not patent eligible. Regarding claim 2 depends on claim 1, thus the rejection of claim 1 is incorporated. “The method according to claim 1, wherein the early-exit neural network includes a plurality of heads, wherein sets of weights for each head are characterized by a posterior distribution of sets of weights and an order of the plurality of heads is given by their position within the early-exit neural network, and wherein the head is preceded by at least one other head in the order of heads, and wherein the likelihood ratio for the possible class or regression result is determined by multiplying a likelihood ratio for the class or regression result determined for the other head with the likelihood ratio determined for the head.” This limitation recites a mental process as well as a mathematical concept. The limitation recites obtaining a posterior distribution of sets of weights, which corresponds to a mathematical concept and a mental process as disclosed above in claim 1, and the likelihood ratio for the possible class or regression result by performing a multiplication operation, which also corresponds to a mathematical concept and a mental process because a person can perform such multiplication. Thus, claim 2 recites abstract ideas rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 2 is not patent eligible. Regarding claim 3 depends on claim 2, thus the rejection of claim 1 is incorporated. “The method according to claim 2, wherein a likelihood ratio for any head of the plurality of heads is determined according to the formula: R t y =   ∏ l = 1 t p l ( y | x , D ) p ( y | x , W l ) , wherein y is a possible class or regression result, l is an index of a head in the plurality of heads, pl is the likelihood of the predictive posterior distribution of the l-th head given a training dataset D and an input x to the EENN, and p is the likehood of predicting y at head l when using the set of sampled weights Wl”. This limitation recites a mental process as well as a mathematical concept. The limitation recites a formula to calculate a likelihood ratio for any head of the plurality of heads within the neural network, wherein such formula corresponds to a mathematical concept, and a person can mentally perform the calculation of the formula. Thus, claim 3 recites abstract ideas rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 3 is not patent eligible. Regarding claim 4 depends on claim 3, thus the rejection of claim 1 is incorporated. “The method according to claim 3, wherein the first element which characterizes a regression result and the confidence interval is determined by determining bounds of the confidence interval, wherein the bounds of the confidence interval are roots of a function characterized by the formula: log ⁡ R t y -   log ⁡ 1 α = 0 , wherein 1/α is the predefined threshold.” This limitation recites a mental process as well as a mathematical concept. The limitation recites a formula to calculate bounds of the confidence interval the neural network, wherein such formula corresponds to a mathematical concept, and a person can mentally perform the calculation of the formula. Thus, claim 4 recites abstract ideas rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 4 is not patent eligible. Regarding claim 6 depends on claim 2, thus the rejection of claim 2 is incorporated. “The method according to claim 2, wherein the head is at a last position within the order of heads” The additional element recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application or amount to significantly more than the judicial exception. The claim merely recites the application of the neural network head, without providing functional details such as architecture or a function of the neural network head or the neural network head perform an unconventional function. Thus, claim 6 recites additional elements rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 6 is not patent eligible. Regarding claim 7 depends on claim 1, thus the rejection of claim 1 is incorporated. “The method according to claim 1, further comprising determining a third element characterizing classification or regression result, wherein the first element is provided as third element when the class characterized by the first element or the regression result characterized by the first element is within the confidence interval characterized by the second element and wherein otherwise a value characterizing a rejection is provided as third element” This limitation recites a mental process as well as a mathematical concept. The limitation recites a mental process of determining a result or rejecting it by comparing the result based on the posterior distribution and the confidence interval to reflect if the result is correct, wherein such determination of whether the result should be accepted or rejected is a mental process, and the calculation using the posterior distribution and the confidence interval are mathematical concepts as disclosed above. Thus, claim 7 recites abstract ideas rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 7 is not patent eligible. Regarding claim 8 depends on claim 1, thus the rejection of claim 1 is incorporated. “The method according to claim 1, wherein the method further includes determining the posterior distribution of the weights using conjugate Bayesian inference” The limitation recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application or amount to significantly more than the judicial exception. The claim merely recites the application of applying conjugate Bayesian inference as a method to obtain posterior distribution, which is a conventional method of machine learning and does not provide any improvement toward the method. Thus, claim 8 recites additional elements rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 8 is not patent eligible. Regarding claim 9 depends on claim 8, thus the rejection of claim 8 is incorporated. “The method according to claim 8, wherein the early-exit neural network is pretrained in a pretraining step, and wherein the determining of the posterior distribution of the weights using conjugate Bayesian inference is conducted after pretraining” The limitation recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application or amount to significantly more than the judicial exception. The claim merely recites the application of applying pre-training machine learning practice and conjugating Bayesian inference as a method to obtain posterior distribution, which are both conventional methods of machine learning. In other words, the claim merely recites applying these techniques for training a neural network and do not provide any improvement toward the neural network algorithm or architecture. Thus, claim 9 recites additional elements rendered at a high level of generality resulting in claims that do not integrate the abstract idea into a practical application or amount to significantly more than the judicial exception. Therefore, claim 9 is not patent eligible. Regarding claim 10, which recites a system, one of the four statutory categories of patentable subject matter, the applicant is directed to the rejection of claim 1 and 9 above. Claim 10 is rejected under the same rationale of claim 1 and 9 because the claim recites similar limitations and processing steps to these claims. Regarding claim 11, which recites a system, one of the four statutory categories of patentable subject matter. Claim 11 recites: “determine a control signal based on the first element and the second element; and control, using the control signal, an actuator and/or a display” The additional element recites an additional element of a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application or amount to significantly more than the judicial exception. The claim merely recites the black-box application of the neural network to obtain results and perform additional control operation using those results without providing any improvement toward the neural network algorithm or unconventional machine learning practice. The applicant is further directed to the rejection of claim 1 above. Claim 11 is rejected under the same rationale of claim 1 because the claim recites similar limitations and processing steps to claim 1. Regarding claim 12, which recites a machine, one of the four statutory categories of patentable subject matter, the applicant is directed to the rejection of claim 1 above. Claim 12 is rejected under the same rationale of claim 1 because the claim recites similar limitations and processing steps to these claims. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-10 are rejected under 35 U.S.C. 103 as being unpatentable over Meronen et.al (NPL: Fixing Overconfidence in Dynamic Neural Networks) in view of Magris et.al (NPL: Bayesian Learning for Neural Networks: an algorithmic survey), further in view of Neiswanger et.al (NPL: Uncertainty quantification using martingales for misspecified Gaussian processes). Regarding claim 1, Meronen teaches a part of the preamble “A computer-implemented method for determining a first element, … wherein the first element characterizes a classification or a regression result of a sensor signal, …, wherein the first element … are determined by an early-exit neural network (EENN)” (page 2 section 1 column 1 “Moreover, we propose aggregating predictions across the exits to utilise uncertainties arising at earlier stages … Our contributions can be summarized as follows: (i) We introduce a probabilistic treatment for early exit dynamic neural networks (DNNs) utilising a Bayesian formulation” Meronen discloses an early-exit neural network having intermediate classifiers that generate prediction results from input image data. Thus, the image data corresponds to the claimed sensor signal, and the prediction produced by an intermediate classifier corresponds to the first element characterizing a classification result determined by the EENN.) Meronen teaches or at least suggests the limitation “determining, by the early-exit neural network, a feature representation of the sensor signal, which is provided to a head of the early-exit neural network” (Page 1 section 1 col.1 “We introduce a probabilistic treatment for early exit dynamic neural networks (DNNs) utilising a Bayesian formulation”, and Page 2 section 2 col. 2 “For a DNN model having nblock intermediate classifiers (see Fig. 1 for a sketch), we refer to the predictive distribution of each of these classifiers as … k = 1,2,..., nblock. The feature representation before the last linear layer of each classifier is referred to as Φi,k = fk(xi) … The prediction pk(ŷi|xi) is obtained from the feature representation …” Meronen discloses an early-exit dynamic neural network having a plurality of intermediate classifiers and determines, for each intermediate classifier, a feature representation before the classifier’s last linear layer, wherein a prediction is obtained from the feature representation. Thus, Meronen’s intermediate classifier corresponds to the claimed head, and the feature representation provided to the intermediate classifier corresponds to the claimed feature representation provided to a head of the early-exit neural network. Meronen further operates on image data. Under the broadest reasonable interpretation consistent with Applicant’s Specification, which expressly states that a sensor signal may be an image recorded by an optical sensor, Meronen’s image input teaches or at least suggests the claimed sensor signal. Accordingly, Meronen teaches or at least suggests determining, by the early-exit neural network, a feature representation of the sensor signal and providing the feature representation to the head of the early-exit neural network, as claimed.) Meronen teaches or at least suggests a part of the limitation “providing a predictive posterior distribution or an argument of a maximum of the predictive posterior distributions as the first element, wherein the predictive posterior distribution is determined based on a posterior distribution of sets of weights of the head …” (Page 3 section 2 col. 2 “Bayesian inference is used to obtain a posterior distribution over the model parameters given the training data … An efficient approximation to the posterior of the model parameters is the Laplace approximation, which performs a second-order Taylor expansion of Eq. (1) around the maximum a posteriori (MAP) estimate of the target distribution, resulting in a Gaussian distribution”, and Page 4 section 3.1 col. 1-2 “To save on computational costs of sampling from a larger dimensional Gaussian on the last-layer parameters N(θMAP,H−1), where θMAP = {WMAP,bMAP} are the last-layer MAP parameters of one exit, we linearly project the Gaussian to a predictive distribution … Absorbing the last layer biases into the weight matrix allows us to shift most of the computation required for sampling to be performed before test time” Meronen discloses using Bayesian inference to obtain a posterior distribution over model parameters and applying a last-layer Laplace approximation at each intermediate exit. Meronen identifies the last-layer parameters of an exit as including a weight matrix and bias and describes a Gaussian distribution over those parameters, which is projected to a predictive distribution for the corresponding intermediate classifier. Thus, Meronen teaches or at least suggests a posterior distribution over sets of weights associated with a head and a predictive posterior distribution determined based on that posterior weight distribution, as claimed. Because the claimed first element may be the predictive posterior distribution itself, Meronen’s provision of the predictive distribution teaches or at least suggests providing the predictive posterior distribution as the first element.) Meronen does not teach a part of the limitation “… the predictive posterior distribution is determined based on … a likelihood of the feature representation given a set of weights of the head”. However, Meronen in view of Magris teaches or at least suggests this part of the limitation (Page 6 section 1.4 example 1 “Equation (2) involves all the ingredients required for performing Bayesian inference on ML models, and specifically neural networks. In the first place, eq. (2) involves a likelihood function for the data Dy conditional on the observed sample Dx and the parameter vector θ… Example 1 … The likelihood is of a certain form and parametrized over a neural network whose weights are denoted by θ” Magris discloses that Bayesian inference for a neural network employs a likelihood function for output data conditional on observed input data and a parameter vector, and further expressly explains that the neural-network parameters θ are the weights of the neural network. In other words, Magris evaluates the likelihood of a possible neural-network output based on an input processed using a particular set of network weights. In combination with the teaching of early exit dynamic neural network comprising the plurality of classifiers by Meronen, Magris’s observed input corresponds to Meronen’s feature representation supplied to the intermediate classifier and the applicable parameters are the weights of that classifier represented by the weight matrix. Magris, in view of Meronen therefore teaches or at least suggests the claimed likelihood associated with the feature representation when a particular set of weights of the head is used.) Meronen does not teach a part of the limitation “sampling a set of weights from the posterior distribution of sets of weights of the head” However, Meronen in view of Magris teaches or at least suggests this limitation (Page 7 section 1.4 “weights in BNNs are stochastic and represented with distributions. A probability distribution over the weights is learned by updating the prior with the evidence supported by the data … the posterior over the weights is, in general, a truly multivariate distribution where independence among its dimensions generally does not hold … an intuitive MC-related approach for approximating the predictive distribution is that of sampling Ns values from the posterior to create Ns realizations of the neural network, each based on a different parameter sample, which are used to provide predictions. This results in a collection of predictions that approximate the actual predictive distribution” Magris discloses that weights in Bayesian neural networks are stochastic and represented by distributions, that a probability distribution over the weights is learned from the prior and observed evidence, and that the resulting posterior over the weights is generally a multivariate distribution. Magris further discloses a Monte-Carlo approach in which values are sampled from the posterior to create different realizations of the neural network, each realization being based on a different parameter sample and producing a corresponding prediction. Because Magris expressly identifies the parameters as neural-network weights, this teaches or at least suggests sampling a set of weights from a posterior distribution over sets of weights. In combination with the teaching of early exit dynamic neural networks comprising the plurality of classifiers by Meronen, those sampled weights correspond to the weights of each intermediate classifier/head.) Before the effective filling date, it would have been obvious to a person of ordinary skill in the art to combine the teaching of early exit dynamic neural networks (DNNs) utilising a Bayesian formulation by Meronen with the teaching of obtaining likelihood based on Bayesian inference for a set of sampled parameters by Magris. The motivation to do so is referred to in Magris’s disclosure (Page 5 section 1.3 “the Bayesian perspective allows the mixing of expert knowledge with experimental evidence. This is quite relevant in small-sample applications where the amount of collected data is inappropriate for classical statistical tools and results to apply (e.g., inference based on asymptotic theory), yet it nevertheless allows the update of the a-priori belief on the parameters, p(θ), into the posterior. On the other hand, the likelihood term captures the aleatoric uncertainty, that is the intrinsic uncertainty naturally embedded in the data”, and page 7 section 1.4 “In general, we may proceed via MC sampling, in which case the estimation leads to a sample, of arbitrary size, approximating the true posterior. The true posterior remains unknown in its exact form, yet MC enables sampling from it and thus estimating an approximate representation … Indeed, an intuitive MC-related approach for approximating the predictive distribution is that of sampling Ns values from the posterior to create Ns realizations of the neural network, … results in a collection of predictions that approximate the actual predictive distribution”. Magris discloses a Bayesian approach in which prior beliefs regarding neural-network parameters are updated using observed data to obtain a posterior distribution, teaches that the likelihood captures uncertainty inherent in the data, and further teaches sampling parameters from the posterior to create different neural-network realizations whose predictions are used to approximate the predictive distribution. Meronen likewise employs a Bayesian formulation for uncertainty estimation at its intermediate classifiers and uses the resulting uncertainty for confidence-based early-exit decisions. Accordingly, a person of ordinary skill in the art would have been motivated to apply Magris’s Bayesian technique to an individual Meronen intermediate classifier/head because the two references employ compatible Bayesian uncertainty techniques; Magris’s likelihood provides information regarding uncertainty in the data, while sampling the posterior weights of that classifier provides different predictions used to characterize the predictive distribution for that particular head. The resulting predictive and uncertainty information would improve Meronen’s confidence-based early-exit determination by providing additional information for deciding whether the current classifier’s prediction should be output or whether processing should continue to a later classifier.) Meronen/Magris does not teach a part of the preamble “A computer-implemented method for determining … a second element … and the second element characterizes a confidence interval of likely classifications or regression results …”. However, Neiswanger teaches or at least suggests this method (page 1 section 1 “This paper instead presents a frequentist approach to uncertainty quantification for GPs (and hence for BO), which uses martingale techniques to construct a confidence sequence (CS) for f∗”. Neiswanger discloses a confidence sequence Ct and expressly teaches that Ct provides a confidence interval for f*(x). Thus, Ct corresponds to the claimed second element characterizing a confidence interval of likely regression results.) Meronen/Magris does not teach the limitation “determining a likelihood ratio for possible classifications or regression results by dividing a value of the predictive posterior distribution at the feature representation by a likelihood of the feature representation given the sampled set of weights” However, Meronen/Magris in view of Neiswanger teaches or at least suggests this limitation (Page 5 section 3.1 “Let GP0(f) represent the prior “density” at function f, and let GPt(f) represent the posterior “density” at f after observing t datapoints … Define the prior-posterior-ratio for any function f as the following real-valued process … equation (7) … Denote the working likelihood of f by … equation (8) … Then, for any function f, the working posterior GP is given by … equation (9) … Substituting theposterior (9) and likelihood (8) into the definition of the prior-posterior-ratio (7), the latter can be more explicitly written as … equation (10)” Neiswanger discloses a likelihood-ratio framework in which a working likelihood is defined for a particular model/function realization and a mixture likelihood ratio compares likelihood support over possible model realizations with the likelihood associated with a particular realization by division. More specifically, Neiswanger’s ratio expresses an integrated or mixture likelihood quantity relative to the likelihood for a particular candidate model realization. Thus, Neiswanger teaches the operation of determining a likelihood ratio by dividing a Bayesian/mixture likelihood quantity by a likelihood corresponding to a particular realization. In view of the combined teaching by Meronen/Magris, wherein Meronen supplies, for an intermediate classifier corresponding to the claimed head, a predictive posterior distribution over possible outputs based on the feature representation, and Magris supplies the likelihood obtained using the sampled weights at the classifier; Neiswanger therefore teaches or at least suggests applying the disclosed likelihood-ratio division to those corresponding quantities.) Meronen/Magris does not teach the limitation “determining the confidence interval as the possible classes or regression results for which the likelihood ratio is equal to or below a predefined threshold and providing the confidence interval or a value characterizing the width of the confidence interval as the second element” However, Meronen/Magris in view of Neiswanger teaches or at least suggests this limitation (Page 7 section 3.2 “we use Lemma 1 to construct the following confidence sequence and use Ville’s inequality to justify its correctness. Define Ct := {f ∈ Ꞙ :Rt(f)< 1/α} … We claim that f∗ is an element of the confidence set Ct, through all of time, with high probability … Ct is a confidence band for the entire function f∗, meaning that it is uniform over both X and time, meaning that it provides a confidence interval for f∗(x) that is valid simultaneously for all times and for all x (on the grid, for simplicity)”. Neiswanger discloses constructing a confidence set Ct by retaining candidate functions for which the likelihood-related ratio Rt(f) is less than a predefined threshold 1/α, and further explains that the resulting confidence set is a confidence band that provides a confidence interval for the corresponding function value f∗(x). Accordingly, Neiswanger teaches the sequence of determining a ratio for candidate values, comparing the ratio with a predefined threshold, retaining the candidate values satisfying the threshold, and using those retained values to form a confidence interval. The candidate value f(x) corresponds to the claimed regression-result alternative, because it represents a possible numerical regression result. Neiswanger therefore teaches or at least suggests determining the confidence interval from possible regression results whose likelihood ratio satisfies a predefined threshold and providing that confidence interval as the claimed second element.) Before the effective filling date, it would have been obvious to a person of ordinary skill in the art to combine the teaching of early exit dynamic neural networks (DNNs) utilising a Bayesian formulation by Meronen, and the teaching of obtaining likelihood based on Bayesian inference for a set of sampled parameters by Magris with the teaching of obtaining a ratio based on likelihoods and a confidence sequence by Neiswanger. The motivation to do so is referred to in Neiswanger’s disclosure (Page 1 section 1 “This paper instead presents a frequentist approach to uncertainty quantification for GPs (and hence for BO), which uses martingale techniques to construct a confidence sequence (CS) for f∗, irrespective of misspecification of the prior. A CS is a sequence of (data-dependent) sets that are uniformly valid over time … The price of such a robust guarantee is that if the prior was indeed accurate, then our confidence sets are looser than those derived from the posterior”, and page 7 section 3.2 “we use Lemma 1 to construct the following confidence sequence and use Ville’s inequality to justify its correctness. Define Ct := {f ∈ Ꞙ :Rt(f)< 1/α}”. Neiswanger discloses using a likelihood-based ratio as a martingale to construct a confidence sequence by determining which possible values satisfy a predefined ratio threshold, and further teaches that the resulting confidence sequence provides a confidence interval that remains valid over time even when the prior is misspecified. Because the Meronen/Magris combination provides predictive-posterior and likelihood information for an intermediate classifier/head, a person of ordinary skill in the art would have been motivated to apply Neiswanger’s ratio-based confidence technique to that information to evaluate the possible prediction values and construct a confidence interval for the classifier’s prediction. The resulting confidence interval would provide a defined range of likely results for evaluating the reliability of the current prediction and thereby improve Meronen’s determination of whether to accept the current classifier’s prediction or continue processing to a later classifier.) Regarding claim 2 depends on claim 1, thus the rejection of claim 1 is incorporated. Meronen teaches or at least suggests “The method according to claim 1, wherein the early-exit neural network includes a plurality of heads, wherein sets of weights for each head are characterized by a posterior distribution of sets of weights and an order of the plurality of heads is given by their position within the early-exit neural network, and wherein the head is preceded by at least one other head in the order of heads,…” (Page 2 section 2 col. 2 “For a DNN model having nblock intermediate classifiers (see Fig. 1 for a sketch), we refer to the predictive distribution of each of these classifiers as … k = 1,2,..., nblock. The feature representation before the last linear layer of each classifier is referred to as Φi,k = fk(xi) … The prediction … is obtained from the feature representation”, Page 3 section 2 col. 2 “Bayesian inference is used to obtain a posterior distribution over the model parameters given the training data … An efficient approximation to the posterior of the model parameters is the Laplace approximation, which performs a second-order Taylor expansion of Eq. (1) around the maximum a posteriori (MAP) estimate of the target distribution, resulting in a Gaussian distribution”, and Page 4-5 section 3.2 “The predictions from different intermediate classifiers are neither independent nor equal, as they are predictions from different stages of the same computational pipeline, and later classifiers have more capacity compared to earlier classifiers” Meronen discloses an early-exit DNN having a plurality of intermediate classifiers, wherein the predictive distribution of the classifiers is indexed k = 1,2,..., nblock. The intermediate classifiers correspond to the claimed heads. As discussed above with respect to claim 1, Meronen employs Bayesian inference and a Laplace approximation of the model parameters, including the last-layer weights associated with an intermediate classifier, thereby characterizing the weights by a posterior distribution. Meronen further explains that the intermediate classifiers provide predictions from different stages of the same computational pipeline and expressly distinguishes earlier classifiers from later classifiers. Accordingly, the successive positions of Meronen’s intermediate classifiers within the common computational pipeline provide an order of the corresponding heads, and a later intermediate classifier necessarily has at least one earlier intermediate classifier preceding it. Thus, Meronen teaches or at least suggests a plurality of ordered heads, posterior-distributed weights associated with the heads, and a head preceded by at least one other head in the order, as claimed.) Meronen in view of Neiswanger teaches or at least suggests “… wherein the likelihood ratio for the possible class or regression result is determined by multiplying a likelihood ratio for the class or regression result determined for the other head with the likelihood ratio determined for the head” (Neiswanger discloses at Page 3 section 2 “At time t (we switch from n to t to emphasize temporality), assume we have already evaluated f∗ at points { X i }   i = 1 t - 1 and obtained observations { Y i }   i = 1 t - 1 . To determine the next domain point Xt …”, Page 4 section 2 “To make the sequential aspect of BO explicit … denote the sigma-field of the first t observations, which captures the information known at time t … Since Dt ⊃ Dt−1 … Using this language, the acquisition function φt is then predictable … meaning that it is measurable with respect to Dt−1 and is hence determined with only the data available after t − 1 steps”, and Page 6 section 3.1 Lemma 1 “Fix any arbitrary f∗∈F, and assume data-generating model (2). Choose any acquisition function φt, any working prior GP0 and construct the working posterior GPt … To conclude the proof, we just need to argue that the braced term in the last expression equals one as claimed. This term can be recognized as integrating a likelihood ratio, which equals one”. Neiswanger discloses the temporal/sequential nature of its process by explaining that at time t, information through t-1 has already been obtained and that Dt contains the information of Dt-1. More particularly, in the proof of Lemma 1, equality (i), Neiswanger rewrites the likelihood-ratio calculation at the current step using the likelihood-ratio quantity accumulated through t-1 multiplied by an additional term associated with the observation at step t; Neiswanger immediately explains that this additional term is recognized as a likelihood ratio. Thus, although Neiswanger’s t-1 and t concern successive observations rather than neural-network heads, Neiswanger teaches the underlying operation of carrying forward a likelihood ratio based on preceding information and multiplying it by a likelihood-ratio contribution associated with the current information. In view of Meronen’s ordered intermediate classifiers corresponding to the claimed heads, a person of ordinary skill would have recognized that this same sequential likelihood-ratio accumulation may be applied across the intermediate classifiers, which teaches or at least suggests the likelihood ratio determined for a preceding head is multiplied by the likelihood ratio determined for the current head, as claimed.) Regarding claim 3 depends on claim 2, thus the rejection of claim 1 is incorporated. Meronen in view of Magris and Neiswanger teaches or at least suggests “The method according to claim 2, wherein a likelihood ratio for any head of the plurality of heads is determined according to the formula: R t y =   ∏ l = 1 t p l ( y | x , D ) p ( y | x , W l ) , wherein y is a possible class or regression result, l is an index of a head in the plurality of heads, pl is the likelihood of the predictive posterior distribution of the l-th head given a training dataset D and an input x to the EENN, and p is the likehood of predicting y at head l when using the set of sampled weights Wl” (Neiswanger discloses at Page 5 section 3.1 “Let GP0(f) represent the prior “density” at function f, and let GPt(f) represent the posterior “density” at f after observing t datapoints … Define the prior-posterior-ratio for any function f as the following real-valued process … equation (7) … Denote the working likelihood of f by … equation (8) … Then, for any function f, the working posterior GP is given by … equation (9) … Substituting the posterior (9) and likelihood (8) into the definition of the prior-posterior-ratio (7), the latter can be more explicitly written as … equation (10)”, Meronen discloses at Page 2 section 2 “we refer to the predictive distribution of each of these classifiers as pk(ŷi|xi), k = 1, 2,..., nblock … The prediction pk(ŷi|xi) is obtained from the feature representation”, and Magris discloses at Page 6 section 1.4 “The forward pass parses the input into predictions via some parameter values, such outputs … follow a prescribed likelihood function ... An underlying neural network is implicit in the likelihood term p(Dy|Dx,θ))”. Neiswanger discloses the working likelihood Lt(f) in Eq. (8) using the product operator ∏, thereby multiplying the individual likelihood contributions from indexed stages 1 through t, and subsequently uses that accumulated likelihood in the likelihood-ratio expression of Eq. (10). Meronen, in turn, expressly indexes its intermediate classifiers as k=1,2,…, nblock; because those intermediate classifiers correspond to the claimed heads, Meronen’s index k corresponds to the claimed index l. Meronen’s predicted output y corresponds to the claimed possible class y, its input xi corresponds to the claimed input x, and its training data corresponds to D. Meronen further provides a predictive distribution for each indexed intermediate classifier pk; a person of ordinary skill would have recognized the value assigned by that predictive distribution to a particular possible result y as corresponding to the head-specific predictive-posterior quantity pl, as claimed. Magris discloses that neural-network outputs follow a prescribed likelihood function conditional on the input and neural-network parameters, expressly identifies those parameters as neural-network weights, and teaches posterior sampling of such parameters. Accordingly, Magris’s sampled parameter set corresponds to Wl, and its conditional likelihood term p corresponds to the likelihood p of predicting y using the sampled weights Wl, as claimed. Applying Neiswanger’s disclosed product operation to these corresponding head-specific numerator and denominator quantities therefore teaches or at least suggests forming the ratio Rt(y) as the product of the respective ratios over l-th heads, as claimed.) Regarding claim 4 depends on claim 3, thus the rejection of claim 1 is incorporated. Neiswanger teaches or at least suggests “The method according to claim 3, wherein the first element which characterizes a regression result and the confidence interval is determined by determining bounds of the confidence interval, wherein the bounds of the confidence interval are roots of a function characterized by the formula: log ⁡ R t y -   log ⁡ 1 α = 0 , wherein 1/α is the predefined threshold.” (page 7 section 3.2 “we use Lemma 1 to construct the following confidence sequence and use Ville’s inequality to justify its correctness. Define Ct := {f ∈ Ꞙ :Rt(f)< 1/α}…”, and page 10 section 4.2-(3) “Compute the confidence sequence. We can then use the confidence sequence … Thus we know that Ct is an ellipsoid defined by the superlevel set … To compute Ct, we can traverse outwards from the posterior-prior ratio mean µc until we have found the Mahalanobis distance k to the isocontour … We can therefore view Ct as the k-sigma ellipsoid of the posterior GP (normal distribution) … Using this confidence ellipsoid over f, we can compute a lower confidence bound for the value of f(X’)”. Neiswanger discloses that the confidence region Ct is an ellipsoid (a bounded confidence region) and teaches computing that region by starting from the posterior mean and traversing outward until reaching an isocontour (the boundary defined by the likelihood-ratio threshold 1/α). Meronen discloses that the intermediate classifiers generate prediction results, corresponding to the claimed classification or regression results. For a one-dimensional regression result, a person of ordinary skill in the art would have understood that the ellipsoidal confidence region reduces to an interval, such that traversing outward in the two directions identifies the lower and upper bounds where the likelihood ratio reaches 1/α. A person of ordinary skill would further have recognized that, because the likelihood ratio is positive, taking the logarithm of the boundary relationship and rearranging the resulting terms would have obtained an equation that suggests the claimed logarithmic equation for determining the lower and upper bounds. Thus, the resulting solutions correspond to the claimed roots defining the confidence-interval bounds.) Regarding claim 5 depends on claim 2, thus the rejection of claim 2 is incorporated. Meronen in view of Neiswanger teaches or at least suggests “The method according to claim 2, wherein the head is not a last head in the plurality of heads and wherein the predictive posterior distribution or a maximum of the predictive posterior distribution is provided as first element and the confidence interval or a value characterizing a width of the confidence interval is provided as the second element when the confidence interval is smaller than or equal to a predefined threshold but not empty and wherein otherwise the predictive posterior distribution or a maximum of the predictive posterior distribution corresponding to a head following the head is provided as first element and a second confidence interval or a value characterizing the width of the second confidence interval corresponding to the head following the head is provided as confidence interval” (Meronen discloses at page 4-5 section 3.2 “The predictions from different intermediate classifiers are neither independent nor equal, as they are predictions from different stages of the same computational pipeline, and later classifiers have more capacity compared to earlier classifiers … the uncertainty should be high for the model to be able to recognize these samples as ‘hard’, and continue their evaluation to the next block”, page 5 section 4 “Early exiting decisions were based on model predicted confidence, for which thresholds tk,k = 1,2,...,nblock were calculated on the validation set”, and Page 3 section 2 col. 2 “Bayesian inference is used to obtain a posterior distribution over the model parameters given the training data”. Neiswanger discloses at Page 7 section 3.2 “we use Lemma 1 to construct the following confidence sequence and use Ville’s inequality to justify its correctness. Define Ct := {f ∈ Ꞙ :Rt(f)< 1/α} … We claim that f∗ is an element of the confidence set Ct, through all of time, with high probability … let |Ct| denote its size … Intuitively, if the working prior GP0 was accurate, … then |Ct| will be (relatively) small … If the working prior GP0 was inaccurate, … then |Ct| will be (relatively) large … Ct is a confidence band for the entire function f∗, meaning that it is uniform over both X and time, meaning that it provides a confidence interval for f∗(x) that is valid simultaneously for all times and for all x”. Meronen discloses a plurality of intermediate classifiers positioned at different stages of the same computational pipeline and teaches that uncertain samples continue evaluation to a following block, while early-exit decisions are based on model-predicted confidence using respective thresholds. Because the intermediate classifiers correspond to the claimed heads, a current intermediate classifier corresponds to a head that is not the last head and a later intermediate classifier corresponds to a head following the current head. Meronen further provides a predictive distribution for each intermediate classifier. Neiswanger discloses a confidence set Ct, characterizes | Ct | as the size of the confidence set, explains that such size may be relatively small or large, and teaches that Ct provides a confidence interval. Accordingly, a person of ordinary skill in the art would have recognized using the size of Neiswanger’s confidence interval as the confidence criterion in Meronen’s threshold-based early-exit framework, such that when the confidence interval is sufficiently small and nonempty, the current-head predictive posterior distribution and confidence interval are provided, whereas otherwise processing continues to the following head, for which the corresponding predictive posterior distribution and subsequent confidence interval are provided as claimed.) Regarding claim 6 depends on claim 2, thus the rejection of claim 2 is incorporated. Meronen teaches or at least suggests “The method according to claim 2, wherein the head is at a last position within the order of heads” (Page 4-5 section 3.2 “The predictions from different intermediate classifiers are neither independent nor equal, as they are predictions from different stages of the same computational pipeline, and later classifiers have more capacity compared to earlier classifiers)” Meronen discloses intermediate classifiers positioned at different stages of the same computational pipeline, including earlier and later classifiers. In view of the ordered plurality of intermediate classifiers previously mapped to the claimed plurality of heads in claim 2, a person of ordinary skill in the art would have understood that the final classifier in that ordered sequence corresponds to a head at the last position within the order of heads, as claimed.) Regarding claim 7 depends on claim 1, thus the rejection of claim 1 is incorporated. Meronen in view of Neiswanger teaches or at least suggests “The method according to claim 1, further comprising determining a third element characterizing classification or regression result, wherein the first element is provided as third element when the class characterized by the first element or the regression result characterized by the first element is within the confidence interval characterized by the second element and wherein otherwise a value characterizing a rejection is provided as third element” (Meronen discloses at page 4 section 3 “we leverage the early exit structure of the DNN architecture and propose an approach that uses Laplace approximations and model-internal ensembling. The motivation for improving uncertainty estimation is that if the intermediate classifiers in the DNN can more accurately estimate the uncertainty in their predictions, they can better recognise which samples are hard and require further computation, and for which samples the prediction is already confident enough to be exited early”, and Neiswanger discloses at page 7 section 3.2 “we use Lemma 1 to construct the following confidence sequence and use Ville’s inequality to justify its correctness. Define Ct := {f ∈ Ꞙ :Rt(f)< 1/α} … We claim that f∗ is an element of the confidence set Ct, through all of time, with high probability … Ct is a confidence band for the entire function f∗, meaning that it is uniform over both X and time, meaning that it provides a confidence interval for f∗(x) that is valid simultaneously for all times and for all x”. Meronen discloses intermediate classifiers that determine prediction results and estimate uncertainty to determine whether a prediction is sufficiently confident to be accepted or instead requires further computation. Neiswanger discloses a confidence set Ct that provides a confidence interval identifying values supported by the confidence determination. Accordingly, a person of ordinary skill in the art would have recognized providing the prediction characterized by the first element as the third element when that prediction is within the confidence interval and, when the prediction is outside the confidence interval and therefore is not accepted, providing a value indicating rejection as the third element. Providing such a rejection value would have been a routine and predictable way of representing the disclosed rejection determination, as claimed.) Regarding claim 8 depends on claim 1, thus the rejection of claim 1 is incorporated. Magris teaches or at least suggests “The method according to claim 1, wherein the method further includes determining the posterior distribution of the weights using conjugate Bayesian inference” (page 6 section 1.4 example 1 “The likelihood is of a certain form and parametrized over a neural network whose weights are denoted by θ, …”, and page 7 section 1.4 “The inference goal is the posterior distribution. (i) If the problem has a form for which the posterior can be solved analytically, we find p(θ|DxDy) to be of a known parametric form and identify the parameters characterizing it (standard Bayesian setting, so-called conjugacy between the prior and the likelihood”. Magris discloses a neural network whose weights are represented by the parameter \theta and teaches that the goal of Bayesian inference is to determine the posterior distribution p(θ|DxDy). Magris further teaches that when the posterior is analytically solvable, the inference is performed in the standard Bayesian setting of conjugacy between the prior and the likelihood. Because θ expressly represents the neural-network weights, Magris therefore teaches or at least suggests determining the posterior distribution of the weights using conjugate Bayesian inference, as claimed.) Regarding claim 9 depends on claim 8, thus the rejection of claim 8 is incorporated. Meronen in view of Magris teaches or at least suggests “The method according to claim 8, wherein the early-exit neural network is pretrained in a pretraining step, and wherein the determining of the posterior distribution of the weights using conjugate Bayesian inference is conducted after pretraining” (page 1-2 section 1 “In this work, we propose a new probabilistic treatment of the multiple exits of DNNs by applying a Bayesian formulation combined with efficient post-hoc approximate inference using multiple last-layer Laplace approximations … works out of the box without retraining”, and page 2 section 2 “Here, a model is trained on the training set Dtrain with an ‘unlimited’ computational budget and tested on a set of test samples” Meronen discloses an early-exit DNN having multiple exits that is first trained on a training set and thereafter applies a Bayesian formulation using post-hoc approximate inference, which operates without retraining. Thus, the initial training of the DNN corresponds to the claimed pretraining step, while the Bayesian posterior inference is performed after such pretraining. In view of the conjugate Bayesian posterior determination incorporated from claim 8, Meronen in combination with Magris teaches or at least suggests determining the posterior distribution of the weights using conjugate Bayesian inference after pretraining, as claimed.) Regarding claim 10, the applicant is directed to the rejection of claim 1 and 9 above. Claim 10 is rejected under the same rationale of claim 1 and 9 because the claim recites similar limitations and processing steps to these claims. Claims 11, 12 is rejected under 35 U.S.C. 103 as being unpatentable over Meronen et.al (NPL: Fixing Overconfidence in Dynamic Neural Networks) in view of Magris et.al (NPL: Bayesian Learning for Neural Networks: an algorithmic survey), further in view of Neiswanger et.al (NPL: Uncertainty quantification using martingales for misspecified Gaussian processes, further in view of Gómez et.al (US 20190332918 A1) Regarding claim 11, Meronen/Magris/Neiswanger does not teach “determine a control signal based on the first element and the second element; and control, using the control signal, an actuator and/or a display”. However, Gómez teaches this limitation (paragraph 18 “the neural network is to predict the value of the state of the target system based on the first measurement and a prior sequence of values of a control signal previously generated to control the target system during the time interval between the prior time and the future time”, and paragraph 26 “uses the measured output(s) and reference value(s) that represent the corresponding desired output(s) of the plant to generate the control signal to be provided to the actuator”. Gómez discloses using neural-network prediction information to generate a control signal and providing the control signal to an actuator for controlling a target system. In view of the first element and second element already provided by the Meronen/Magris/Neiswanger combination, Gómez teaches or at least suggests using the determined predictive information to generate the claimed control signal and controlling an actuator using that control signal, as claimed.) Before the effective filling date, it would have been obvious to a person of ordinary skill in the art to combine the teaching of early exit dynamic neural networks (DNNs) utilising a Bayesian formulation by Meronen, the teaching of obtaining likelihood based on Bayesian inference for a set of sampled parameters by Magris, and the teaching of obtaining a ratio based on likelihoods and a confidence sequence by Neiswanger with the teaching of … by Gómez. The motivation to do so is referred to in Gómez’s disclosure (paragraph 31 “some example predictive feedback control solutions implemented in accordance with the teachings of this disclosure include an example state prediction neural network to prevent performance degradation when a wireless network, which typically has larger latency than a wired network, is used instead of, or is used to replace, a wired network in the feedback control system. As disclosed in further detail below, the state prediction neural network utilizes information provided by the wireless devices ..., to predict a future state of the controlled system, which can be applied to the controller that is to generate the control signal for use by the actuator”. Gómez discloses predictive feedback control in which neural-network prediction information is used to generate a control signal for an actuator and teaches that such predictive control can prevent or reduce performance degradation caused by system delay. Accordingly, a person of ordinary skill in the art would have been motivated to apply Gómez’s predictive-control technique to the prediction and confidence information provided by the Meronen/Magris/Neiswanger combination so that the determined predictive information can be used to generate a control signal for controlling a physical system through an actuator, thereby improving feedback-control performance.) The applicant is further directed to the rejection of claim 1 above. Claim 11 is rejected under the same rationale of claim 1 because the claim recites similar limitations and processing steps to claim 1. Regarding claim 12, Gómez teaches or at least suggests “A non-transitory machine-readable storage medium on which is stored a computer program” (paragraph 86 “In these examples, the machine readable instructions may be one or more executable programs or portion(s) of an executable program for execution by a computer processor, such as the processors … The one or more programs, or portion(s) thereof, may be embodied in software stored on a non-transitory computer readable storage medium”. Gómez discloses executable machine-readable instructions embodied as software and expressly teaches storing such software on a non-transitory computer-readable storage medium for execution by a processor. Thus, Gómez teaches the claimed non-transitory machine-readable storage medium storing a computer program, while the processing steps performed by the computer program are supplied by the Meronen/Magris/Neiswanger combination as mapped above.) The motivation to combine the teachings is similar to the motivation in claim 11. The applicant is further directed to the rejection of claim 1 above. Claim 12 is rejected under the same rationale of claim 1 because the claim recites similar limitations and processing steps to claim 1. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DUY TU DIEP whose telephone number is (703)756-1738. The examiner can normally be reached M-F 8-4:30. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DUY T DIEP/Examiner, Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Sep 03, 2024
Application Filed
Sep 21, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748949
A Central Node and Method Therein for Enabling an Aggregated Machine Learning Model from Local Machine Learnings Models in a Wireless Communications Newtork
4y 1m to grant Granted Sep 29, 2026
Patent 12743603
KNOWLEDGE GRAPH REASONING MODEL, SYSTEM, AND REASONING METHOD BASED ON BAYESIAN FEW-SHOT LEARNING
3y 11m to grant Granted Sep 22, 2026
Patent 12725004
NEURAL ARCHITECTURE SEARCH BASED OPTIMIZED DNN MODEL GENERATION FOR EXECUTION OF TASKS IN ELECTRONIC DEVICE
5y 5m to grant Granted Sep 01, 2026
Patent 12651158
NEURAL NETWORK TRAINING METHOD AND APPARATUS USING TREND
4y 1m to grant Granted Jun 09, 2026
Patent 12608642
MODEL PARAMETER LEARNING METHOD AND MOVEMENT MODE DETERMINATION METHOD
4y 7m to grant Granted Apr 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
37%
Grant Probability
61%
With Interview (+23.7%)
4y 4m (~2y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 35 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month