Prosecution Insights
Last updated: August 17, 2026
Application No. 18/587,965

REGRESSION ESTIMATION DEVICE, REGRESSION ESTIMATION METHOD, PROGRAM, AND METHOD FOR GENERATING TRAINED MODEL

Non-Final OA §101§103
Filed
Feb 27, 2024
Priority
Aug 31, 2021 — JP 2021-141458 +1 more
Examiner
ADMASU, MAHLIET TASEW
Art Unit
Tech Center
Assignee
Fujifilm Holdings Corporation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
14 currently pending
Career history
12
Total Applications
across all art units

Statute-Specific Performance

§101
31.5%
-8.5% vs TC avg
§103
57.4%
+17.4% vs TC avg
§112
9.3%
-30.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103
DETAILED ACTION This communication is in response to the Application No. 18/587,965 filed February 27, 2024 in which Claims 1 - 20 are presented for examination. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 1-20 are rejected under 35 U.S.C. 101 because these claimed inventions are directed to an abstract idea without significantly more. Regarding Claim 1: Step 1: Claim 1 is a device type claim. Therefore, Claims 1-15 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. to input the plurality of data items to […]to estimate a plurality of sets of estimated values and certainties of the estimated values from the plurality of data items (mental process - to input the plurality of data items to […]to estimate a plurality of sets of estimated values and certainties of the estimated values from the plurality of data items may be performed mentally or using pen and paper by a user observing/analyzing the plurality of data items and accordingly using judgment/evaluation and mathematical/statistical reasoning to determine estimated values and certainty values based on said analysis) to integrate estimation results of the plurality of sets on the basis of the plurality of sets of the estimated values and the certainties of the estimated values estimated [...] (mental process – to integrate estimation results of the plurality of sets on the basis of the plurality of sets of the estimated values and the certainties of the estimated values estimated[..] may be performed mentally or using pen and paper by a user observing/analyzing the plurality of estimated values and certainty values and accordingly using judgment/evaluation and mathematical/statistical reasoning to combine the estimation results based on said estimated values and certainty values) Step 2A Prong 2: This judicial exception is not integrated into a practical application. one or more processors (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) and one or more storage devices […](recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, storage devices and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) to receive an input of a plurality of data items (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)) […] a single regression model […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a regression model without significantly more) […]by the regression model (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a regression model without significantly more) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. one or more processors (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) and one or more storage devices […](recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, storage devices and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) to receive an input of a plurality of data items (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) […] a single regression model […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a regression model without significantly more) […]by the regression model (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a regression model without significantly more) For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 1 - 15. The additional limitations of the dependent claims are addressed below. Regarding Claim 2: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 2 depends on. estimate a probability distribution having the estimated value as a random variable on the basis of the estimated value and the certainty of the estimated value (mental process - estimate a probability distribution having the estimated value as a random variable on the basis of the estimated value and the certainty of the estimated value may be performed mentally or using pen and paper by a user observing/analyzing the estimated value and certainty value and accordingly using judgment/evaluation and mathematical/statistical reasoning to determine a probability distribution based on said estimated value and certainty value) integrate the probability distributions of the plurality of sets to generate an integrated distribution (mental process – integrate the probability distributions of the plurality of sets to generate an integrated distribution may be performed mentally or using pen and paper by a user observing/analyzing the probability distributions of the plurality of sets and accordingly using judgment/evaluation and mathematical/statistical reasoning to combine the probability distributions into an integrated distribution based on said analysis) and specify a final estimated value on the basis of the integrated distribution (mental process - specify a final estimated value on the basis of the integrated distribution may be performed mentally or using pen and paper by a user observing/analyzing the integrated distribution and accordingly using judgment/evaluation and mathematical/statistical reasoning to select or determine a final estimated value based on said integrated distribution) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 3: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 3 depends on. estimate a probability distribution having the estimated value as a random variable on the basis of the estimated value and the certainty of the estimated value (mental process - estimate a probability distribution having the estimated value as a random variable on the basis of the estimated value and the certainty of the estimated value may be performed mentally or using pen and paper by a user observing/analyzing the estimated value and the certainty of the estimated value and accordingly using judgment/evaluation and mathematical/statistical reasoning to determine a probability distribution based on said analysis) and specify a value at which a product of probabilities at the same random variable is maximized on the basis of the probability distribution of each of the plurality of sets (mental process - specify a value at which a product of probabilities at the same random variable is maximized on the basis of the probability distribution of each of the plurality of sets may be performed mentally or using pen and paper by a user observing/analyzing the probability distribution of each of the plurality of sets and accordingly using judgment/evaluation and mathematical/statistical reasoning to calculate products of probabilities at the same random variable and select the value at which the product is maximized) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 4: Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 4 depends on. perform variable conversion to convert the estimated value output from the regression model into a first parameter of a probability distribution model (mathematical concept – perform variable conversion to convert the estimated value output from the regression model into a first parameter of a probability distribution model recites a mathematical concept because it involves applying a mathematical function to convert one numerical/model-output value into another numerical parameter value. Paragraph [0042] states “FIG. 3 is a graph of a function y = 1 / l o g ⁡ ( 1 + e x p ⁡ ( - x ) ) used for variable conversion,” showing that the variable conversion is performed using a mathematical function/formula) perform variable conversion to convert a value indicating the certainty output from the regression model into a second parameter of the probability distribution model (mathematical concept – perform variable conversion to convert a value indicating the certainty output from the regression model into a second parameter of the probability distribution model recites a mathematical concept because it involves applying a mathematical function/formula to convert one numerical/model-output value into another numerical parameter value. Paragraph [0042] describes “FIG. 3 is a graph of a function y = 1 / l o g ⁡ ( 1 + e x p ⁡ ( - x ) ) used for variable conversion,” showing that the variable conversion is performed using a mathematical function/formula) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 5: Step 2A Prong 1: See the rejection of Claim 4 above, which Claim 5 depends on. Step 2A Prong 2 & Step 2B: wherein the probability distribution model is a Laplace distribution (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the probability distribution model is a Laplace distribution does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 4. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 6: Step 2A Prong 1: See the rejection of Claim 4 above, which Claim 6 depends on. Step 2A Prong 2 & Step 2B: wherein the probability distribution model is a Gaussian distribution (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the probability distribution model is a Gaussian distribution does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 4. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 7: Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 7 depends on. perform logarithmic conversion to take a logarithm of the probability distribution (mathematical concept – perform logarithmic conversion to take a logarithm of the probability distribution recites a mathematical concept because it involves applying a logarithmic mathematical operation to a probability distribution. Taking a logarithm of a probability distribution is a mathematical calculation/formula-based conversion of one numerical/probability expression into another numerical/logarithmic expression) calculate a sum of logarithmic probability densities corresponding to the probability distributions of the plurality of sets during the integration (mathematical concept – calculate a sum of logarithmic probability densities corresponding to the probability distributions of the plurality of sets during the integration recites a mathematical concept because it involves applying logarithmic and summation operations to probability density values. The limitation requires taking logarithmic probability densities associated with multiple probability distributions and mathematically adding them during the integration process) calculate a value at which a simultaneous logarithmic probability density is maximized (mathematical concept – calculate a value at which a simultaneous logarithmic probability density is maximized recites a mathematical concept because it involves applying mathematical/statistical operations to determine an optimum value. The limitation requires evaluating a logarithmic probability density and mathematically selecting the value that maximizes that density) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 8: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 8 depends on. Step 2A Prong 2 & Step 2B: wherein the regression model includes a trained model generated by performing machine learning using training data in which data for input and a teaching signal are associated with each other (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of training a model/performing machine learning using training data without significantly more) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 9: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 9 depends on. Step 2A Prong 2 & Step 2B: wherein the regression model is configured using a convolutional neural network (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the regression model is configured using a convolutional neural network does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 10: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 10 depends on. Step 2A Prong 2 & Step 2B: wherein the plurality of data items are medical images (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the plurality of data items are medical images does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 11: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 11 depends on. Step 2A Prong 2 & Step 2B: wherein the plurality of data items include different partial images included in a three-dimensional image (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the plurality of data items include different partial images included in a three-dimensional image does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 12: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 12 depends on. Step 2A Prong 2 & Step 2B: wherein the plurality of data items include generated images that are generated on the basis of different partial images included in a three-dimensional image (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the plurality of data items include generated images that are generated on the basis of different partial images included in a three-dimensional image does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 13: Step 2A Prong 1: See the rejection of Claim 10 above, which Claim 13 depends on. Step 2A Prong 2 & Step 2B: wherein the estimated value is an elapsed time from injection of a contrast agent (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the estimated value is an elapsed time from injection of a contrast agent does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 10. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 14: Step 2A Prong 1: See the rejection of Claim 10 above, which Claim 14 depends on. Step 2A Prong 2 & Step 2B: wherein the estimated value is a value that indicates a position of a specific object (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the estimated value is a value that indicates a position of a specific object does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 10. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 15: Step 2A Prong 1: See the rejection of Claim 11 above, which Claim 15 depends on. Step 2A Prong 2 & Step 2B: wherein the estimated value is a value that indicates a position of the partial image in the three-dimensional image (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the estimated value is a value that indicates a position of the partial image in the three-dimensional image does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 11. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 16: Step 1: Claim 16 is a method type claim. Therefore, Claims 16-17 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. inputting the plurality of data items to […]to estimate a plurality of sets of estimated values and certainties of the estimated values from the plurality of data items (mental process - to input the plurality of data items to […]to estimate a plurality of sets of estimated values and certainties of the estimated values from the plurality of data items may be performed mentally or using pen and paper by a user observing/analyzing the plurality of data items and accordingly using judgment/evaluation and mathematical/statistical reasoning to determine estimated values and certainty values based on said analysis) integrating estimation results of the plurality of sets on the basis of the plurality of sets of the estimated values and the certainties of the estimated values estimated [...] (mental process – to integrate estimation results of the plurality of sets on the basis of the plurality of sets of the estimated values and the certainties of the estimated values estimated[..] may be performed mentally or using pen and paper by a user observing/analyzing the plurality of estimated values and certainty values and accordingly using judgment/evaluation and mathematical/statistical reasoning to combine the estimation results based on said estimated values and certainty values) Step 2A Prong 2: This judicial exception is not integrated into a practical application. receiving an input of a plurality of data items (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)) […] a single regression model […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a regression model without significantly more) […]by the regression model (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a regression model without significantly more) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. receiving an input of a plurality of data items (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) […] a single regression model […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a regression model without significantly more) […]by the regression model (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a regression model without significantly more) For the reasons above, Claim 16 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 16 - 17. The additional limitations of the dependent claims are addressed below. Regarding Claim 17: Step 2A Prong 1: See the rejection of Claim 16 above, which Claim 17 depends on. Step 2A Prong 2 & Step 2B: A non-transitory, computer-readable tangible recording medium on which a program for causing, when read by a computer, a processor of the computer to execute the regression estimation method according to claim 16 is recorded (recited at a high-level of generality (i.e., a non-transitory, computer-readable tangible recording medium, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 16. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 18: Step 1: Claim 18 is a method type claim. Therefore, Claims 18-20 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. using training data in which data for input and a teaching signal are associated with each other, inputting the data for input to […], and obtaining an output of the estimated value and a value indicating the certainty of the estimated value […] (mental process - using training data in which data for input and a teaching signal are associated with each other, inputting the data for input to […], and obtaining an output of the estimated value and a value indicating the certainty of the estimated value […] may be performed mentally or using pen and paper by a user observing/analyzing the input data and associated teaching signal and accordingly using judgment/evaluation and mathematical/statistical reasoning to determine an estimated value and certainty value based on said analysis) perform variable conversion to convert the estimated value output from the learning model into a first parameter of a probability distribution model (mathematical concept – perform variable conversion to convert the estimated value output from the regression model into a first parameter of a probability distribution model recites a mathematical concept because it involves applying a mathematical function to convert one numerical/model-output value into another numerical parameter value. Paragraph [0042] states “FIG. 3 is a graph of a function y = 1 / l o g ⁡ ( 1 + e x p ⁡ ( - x ) ) used for variable conversion,” showing that the variable conversion is performed using a mathematical function/formula) perform variable conversion to convert a value indicating the certainty output from the learning model into a second parameter of the probability distribution model (mathematical concept – perform variable conversion to convert a value indicating the certainty output from the regression model into a second parameter of the probability distribution model recites a mathematical concept because it involves applying a mathematical function/formula to convert one numerical/model-output value into another numerical parameter value. Paragraph [0042] describes “FIG. 3 is a graph of a function y = 1 / l o g ⁡ ( 1 + e x p ⁡ ( - x ) ) used for variable conversion,” showing that the variable conversion is performed using a mathematical function/formula) Step 2A Prong 2: This judicial exception is not integrated into a practical application. […] a learning model […] from the learning model (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a learning model without significantly more) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. […] a learning model […] from the learning model (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a learning model without significantly more For the reasons above, Claim 18 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 18 - 20. The additional limitations of the dependent claims are addressed below. Regarding Claim 19: Step 2A Prong 1: See the rejection of Claim 18 above, which Claim 19 depends on. Step 2A Prong 2 & Step 2B: wherein the probability distribution model is a Laplace distribution, and in a case where the first parameter is μ, the second parameter is b, and the teaching signal is t, the following expression is used as the loss function: logb+(❘t-μ❘)⁄(b.) (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the probability distribution model is a Laplace distribution and a mathematical expression is used as the loss function does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 18. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 20: Step 2A Prong 1: See the rejection of Claim 18 above, which Claim 20 depends on. Step 2A Prong 2 & Step 2B: wherein the probability distribution model is a Gaussian distribution, and in a case where the first parameter is μ, the second parameter is σ2, and the teaching signal is t, the following expression is used as the loss function: logσ^2+(t-μ)^2⁄2 σ^2 (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the probability distribution model is a Gaussian distribution and a mathematical expression is used as the loss function does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 18. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 8, 10-12, and 14-17are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (hereafter Wang, a non-patent literature reference titled “Aleatoric uncertainty estimation with test-time augmentation for medical image segmentation with convolutional neural networks”) in view of Fukayama et al. (hereinafter Fukayama) (US 10614830). Regarding Claim 1, Wang teaches: to receive an input of a plurality of data items (Wang, Page 2 – Section 2.3, “The transformations for augmentation typically include flipping, cropping, rotating, and scaling training images. Abdulkadir et al. [6] and Ronneberger et al. [31] also used elastic deformations for biomedical image segmentation. Several studies have empirically found that combining predictions of multiple transformed versions of a test image helps to improve the performance”, & Page 5 – Section 4.1.1, “We collected clinical T2-weighted MRI scans of 60 fetuses in the second trimester with SSFSE on a 1.5 Tesla MR system (Aera, Siemens, Erlangen, Germany). The data for each fetus contained three stacks of 2D slices acquired in axial, sagittal and coronal views respectively, with pixel size 0.63–1.58 mm and slice thick- ness 3–6 mm. The gestational age ranged from 19 weeks to 33 weeks. We used 2640 slices from 120 stacks of 40 patients for training, 278 slices from 12 stacks of 4 patients for validation and 1180 slices from 48 stacks of 16 patients for testing”, thus receiving an input of a plurality of data items is disclosed, because Wang teaches using multiple transformed versions of a test image and multiple 2D MRI slices. The transformed test images and MRI slices correspond to the plurality of data items, and Wang’s process of using those images for prediction corresponds to receiving an input of the plurality of data items), to input the plurality of data items to a single regression model […] from the plurality of data items (Wang, Page 2-3 – Section 2.3, “Several studies have empirically found that combining predictions of multiple transformed versions of a test image helps to improve the performance. For example, Matsunaga et al. [17] geometrically transformed test images for skin lesion classification. [32] used a single model to predict multiple transformed copies of unlabeled images for data distillation. Jin et al. [18] tested on samples extended by rotation and translation for pulmonary nodule detection”, & Page 3 – Section 3.2, “In the context of deep learning, let f ( ·) be the function represented by a neural network, and θ represent the parameters learned from a set of training images with their corresponding annotations”, thus to input the plurality of data items to a single regression model […] from the plurality of data items is disclosed, because Wang teaches using multiple transformed versions/copies of a test image and using a single model / neural network function f(·) to make predictions. The multiple transformed test images correspond to the plurality of data items, and Wang’s single model/neural network corresponds to the single regression model that receives those data items and generates prediction results from them), and to integrate estimation results of […] on the basis of […] estimated by the regression model (Wang, Page 3 – Section 3.2, “Then we obtain one possible hidden image with βn and e n based on Eq. (2) , and feed it into the trained network to get its prediction, which is transformed with βn to obtain y n according to Eq. (4) . With the set Y = { y 1 , y 2 , ..., y N } sampled from p ( Y | X ), E ( Y | X ) is estimated as the average of Yand we use it as the final prediction ˆ Y for X : ˆ Y = E(Y | X ) ≈1 N N n =1 y n (8) For classification or segmentation problems, p ( Y | X ) is a discretized distribution. We obtain the final prediction for X by maximum likelihood estimation: ˆ Y = arg max y p(y | X ) ≈Mode Y”, thus and to integrate estimation results of […] on the basis of […] estimated by the regression model is disclosed, because Wang teaches feeding each hidden/transformed image into the trained network to obtain predictions y₁, y₂, …, yN, and then estimating the final prediction as the average of Y. Wang’s predictions y₁, y₂, …, yN correspond to the estimation results, and Wang’s averaging of those predictions corresponds to integrating the estimation results based on the values estimated by the regression model) Wang does not explicitly disclose one or more processors, one or more storage devices that store a program to be executed by the one or more processors, […] to estimate a plurality of sets of estimated values and certainties of the estimated values […] and […] the plurality of sets […] the plurality of sets of the estimated values and the certainties of the estimated values […]. However, Fukayama teaches: one or more processors (Fukayama, Background - Par. [0017], “Further, the estimator configuring section and the estimating section may be each comprised of a plurality of processors”, thus one or more processors is disclosed) one or more storage devices that store a program to be executed by the one or more processors (Fukayama, Background - Par. [0017], “Further, the estimator configuring section and the estimating section may be each comprised of a plurality of processors and a plurality of memories”, thus one or more storage devices that store a program is disclosed) […] to estimate a plurality of sets of estimated values and certainties of the estimated values […] (Fukayama, Abstract, “The estimating section calculates weights to be added the estimation results output from the regression models, based on the degrees of confidence with respect to the inputs into the regression models.”, & Background - Par. [0012], “The plurality of regression models are each capable of obtaining a probability distribution of the estimation results and a degree of confidence”, thus […] to estimate a plurality of sets of estimated values and certainties of the estimated values […] is disclosed, because Fukayama teaches regression models that output estimation results and degrees of confidence. Fukayama’s estimation results correspond to the estimated values, and Fukayama’s degrees of confidence/probability distributions correspond to the certainties of the estimated values) […] the plurality of sets […] the plurality of sets of the estimated values and the certainties of the estimated values […] (Fukayama, Description - Par. [0022], “The estimating section 4 receives an observation signal to be analyzed as an input, and obtains a mean of the estimation results (target values) and a variance in probability distributions using the regression models 21 to 2 n. Then, weights to be added to the estimation values are calculated based on the degrees of confidence (the degree of confidence is an inverse number of the variance) that are obtained by the regression models 21 to 2 n. Aggregation of weighted estimation results is performed by summing up the weighted estimated results or calculating a weighted sum”, & Description - Par. [0088], “Finally, the estimated values obtained from the individual regression models were aggregated, based on the degrees of confidence, to obtain the VA values. The estimated values were normalized such that the sum of the values should be one (1) to obtain a weight that is proportional to an inverse number of the variance. The weighted sum of the weighted estimated values respectively obtained from the regression models was calculated. The value of the weight sum was determined as the estimation result”, thus […] the plurality of sets […] the plurality of sets of the estimated values and the certainties of the estimated values […] is disclosed, because Fukayama teaches obtaining estimated values and variances, calculating degrees of confidence from the variances, and aggregating the weighted estimation results. Fukayama’s estimated values/target values correspond to the estimated values, Fukayama’s degrees of confidence/inverse variance correspond to the certainties of the estimated values, and Fukayama’s weighted sum of the estimated results corresponds to integrating the plurality of sets based on those values and certainties) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Wang with Fukayama by modifying Wang’s single model test-time augmentation prediction process to use Fukayama’s confidence based regression estimation and aggregation. Wang teaches receiving multiple image data items, inputting transformed image copies/MRI slices to a single neural network model, obtaining multiple predictions y₁, y₂, …, yN, and integrating those predictions by averaging. Fukayama teaches a computer implemented regression estimation system having processors and memories, and further teaches obtaining estimated values with degrees of confidence and aggregating the estimated values based on those confidence values. Therefore, a POSITA would have been motivated to apply Fukayama’s confidence based weighting/aggregation to Wang’s multiple prediction results so that Wang’s final prediction would account for the reliability of each estimated value instead of treating all predictions equally, thereby improving estimation performance (Fukayama, Background - Par. [0018], “The kind of an observation signal is arbitrary. In the music emotion recognition, the observation signal is a music audio signal, and the target value for the observation signal is a music emotion value. It has been confirmed that estimation performance can be improved more than ever by calculating the estimation results (mean values) and degrees of confidence using the regression models in respect of the input music audio signal and aggregating the estimation results through the maximum likelihood estimation using the calculated estimation results and degrees of confidence”) Regarding Claim 2, Wang combined with Fukayama teaches all the limitations of claim 1 as cited above and Fukayama further teaches: estimate a probability distribution having the estimated value as a random variable on the basis of the estimated value and the certainty of the estimated value (Fukayama, Description – Par. [0019], “The distribution of estimation results can be represented as a Gaussian distribution and an inverse number of a variance in probability distributions can be interpreted as the degree of confidence for the estimation results”, & Description – Par. [0022], “The estimating section 4 receives an observation signal to be analyzed as an input, and obtains a mean of the estimation results (target values) and a variance in probability distributions using the regression models 21 to 2 n.”, thus estimate a probability distribution having the estimated value as a random variable on the basis of the estimated value and the certainty of the estimated value is disclosed, because Fukayama teaches representing estimation results as a Gaussian probability distribution using a mean and variance. Fukayama’s mean of the estimation results/target values corresponds to the estimated value, Fukayama’s variance/inverse variance degree of confidence corresponds to the certainty of the estimated value, and Fukayama’s Gaussian distribution corresponds to the probability distribution having the estimated value as a random variable), integrate the probability distributions of the plurality of sets to generate an integrated distribution (Fukayama, Description – Par. [0060], “Here, how to aggregate the estimation results obtained from N different features is discussed. Assuming that each of the estimation errors of εn (n=1, . .. N) follows the Gaussian distribution of zero mean and variance σ2 n, N probability distributions are obtained for the estimated value y as follows. PNG media_image1.png 48 367 media_image1.png Greyscale ”, & Descrption – Par. [0061], “If the estimated values are independent to each other, the joint probability PJ(y), from which N estimation results are obtained, is obtained by calculating a product of the respective probabilities for n where n=1 to N. The joint probability PJ(y) can be obtained by the following expression PNG media_image2.png 112 523 media_image2.png Greyscale ”, thus integrate the probability distributions of the plurality of sets to generate an integrated distribution is disclosed, because Fukayama teaches obtaining N probability distributions for the estimated value and calculating a joint probability PJ(y) by multiplying the respective probabilities. Fukayama’s N probability distributions correspond to the probability distributions of the plurality of sets, and Fukayama’s joint probability PJ(y) corresponds to the integrated distribution generated by integrating the probability distributions), specify a final estimated value on the basis of the integrated distribution (Fukayama, Description – Par. [0063], “A value y which maximizes the joint probability is a value of maximum likelihood estimation with respect to y. To maximize the joint probability PJ(y) for y in the above expression, ξ may be maximized with respect of y. Therefore, the following expression can be obtained by solving dξ2/dy=0 PNG media_image3.png 61 427 media_image3.png Greyscale ”, & Description – Par. [0065], “For example, FIG. 15 illustrates that a weighted mean is calculated in aggregating two estimation results. In an example of FIG. 15, one of the estimation results has an estimated value of −0.3 and a variance of 0.08, and the other estimation result has an estimated value of 0.4 and a variance of 0.2, and then an aggregation result of −0.1, which is the maximum likelihood estimation value, is obtained”, thus specify a final estimated value on the basis of the integrated distribution is disclosed, because Fukayama teaches selecting the value y that maximizes the joint probability PJ(y) and identifying that value as the maximum likelihood estimation value. Fukayama’s joint probability PJ(y) corresponds to the integrated distribution, and Fukayama’s maximum likelihood estimation value corresponds to the final estimated value specified based on that integrated distribution) Regarding Claim 3, Wang combined with Fukayama teaches all the limitations of claim 1 as cited above and Fukayama further teaches: estimate a probability distribution having the estimated value as a random variable on the basis of the estimated value and the certainty of the estimated value (Fukayama, Description – Par. [0019], “The distribution of estimation results can be represented as a Gaussian distribution and an inverse number of a variance in probability distributions can be interpreted as the degree of confidence for the estimation results”, & Description – Par. [0022], “The estimating section 4 receives an observation signal to be analyzed as an input, and obtains a mean of the estimation results (target values) and a variance in probability distributions using the regression models 21 to 2 n.”, thus estimate a probability distribution having the estimated value as a random variable on the basis of the estimated value and the certainty of the estimated value is disclosed, because Fukayama teaches representing estimation results as a Gaussian probability distribution using a mean and variance. Fukayama’s mean of the estimation results/target values corresponds to the estimated value, Fukayama’s variance/inverse variance degree of confidence corresponds to the certainty of the estimated value, and Fukayama’s Gaussian distribution corresponds to the probability distribution having the estimated value as a random variable), specify a value at which a product of probabilities at the same random variable is maximized on the basis of the probability distribution of each of the plurality of sets (Fukayama, Description – Par. [0060], “N probability distributions are obtained for the estimated value y as follows”, & Par. [0061], “If the estimated values are independent to each other, the joint probability PJ(y), from which N estimation results are obtained, is obtained by calculating a product of the respective probabilities for n where n=1 to N”, & Par. [0063], “A value y which maximizes the joint probability is a value of maximum likelihood estimation with respect to y”, thus specify a value at which a product of probabilities at the same random variable is maximized on the basis of the probability distribution of each of the plurality of sets is disclosed, because Fukayama teaches obtaining N probability distributions for the estimated value y, calculating a joint probability PJ(y) by multiplying the respective probabilities, and selecting the value y that maximizes the joint probability. Fukayama’s N probability distributions correspond to the probability distribution of each of the plurality of sets, Fukayama’s product of the respective probabilities corresponds to the product of probabilities at the same random variable, and Fukayama’s value y that maximizes the joint probability corresponds to the specified value) Regarding Claim 8, Wang combined with Fukayama teaches all the limitations of claim 1 as cited above and Wang further teaches: wherein the regression model includes a trained model generated by performing machine learning using training data in which data for input and a teaching signal are associated with each other (Wang, Page 3 – Section 3.2, “In the context of deep learning, let f ( ⋅ ) be the function represented by a neural network, and θ represent the parameters learned from a set of training images with their corresponding annotations”, thus wherein the regression model includes a trained model generated by performing machine learning using training data in which data for input and a teaching signal are associated with each other is disclosed, because Wang teaches a neural network model having parameters learned from training images and their corresponding annotations. Wang’s neural network function f ( ⋅ ) corresponds to the regression model/trained model, Wang’s training images correspond to the data for input, and Wang’s corresponding annotations correspond to the teaching signal associated with the input data) Regarding Claim 10, Wang combined with Fukayama teaches all the limitations of claim 1 as cited above and Wang further teaches: wherein the plurality of data items are medical images (Wang, Page 10 – Section 5, “Experiments with 2D and 3D medical image segmentation tasks showed that uncertainty estimation with our formulated TTA helps to reduce overconfident incorrect predictions encountered by model-based uncertainty es- timation and TTA leads to higher segmentation accuracy than a single-prediction baseline and multiple predictions using test-time dropout”, thus wherein the plurality of data items are medical images is disclosed, because Wang teaches experiments involving 2D and 3D medical image segmentation tasks. Wang’s 2D and 3D images used for medical image segmentation correspond to the plurality of data items, and the medical segmentation context confirms that those data items are medical images) Regarding Claim 11, Wang combined with Fukayama teaches all the limitations of claim 1 as cited above and Wang further teaches: wherein the plurality of data items include different partial images included in a three-dimensional image (Wang, Page 7 – Section 4.2.1, “we used the BraTS 2017 3 [44] training dataset that consisted of volumetric images from 285 studies, with ground truth provided by the organizers. We randomly selected 20 studies for validation and 50 studies for testing, and used the remaining for training. For each study, there were four scans of T1w, T1wce, T2w and FLAIR images, and they had been co-registered”, & Page 7 – Section 4.2.1, “W-Net is a 2.5D network, and we compared using W-Net only in axial view and a fusion of axial, sagittal and coronal views. These two implementations are referred to as W- Net(A) and W-Net(ASC) respectively”, thus wherein the plurality of data items include different partial images included in a three-dimensional image is disclosed, because Wang teaches using volumetric images from the BraTS 2017 dataset and further teaches processing/fusing axial, sagittal, and coronal views. Wang’s volumetric images correspond to the three-dimensional image, and Wang’s axial, sagittal, and coronal views correspond to the different partial images included in the three-dimensional image) Regarding Claim 12, Wang combined with Fukayama teaches all the limitations of claim 1 as cited above and Wang further teaches: wherein the plurality of data items include generated images that are generated on the basis of different partial images included in a three-dimensional image (Wang, Page 7 – Section 4.2.1, “we used the BraTS 2017 3 [44] training dataset that consisted of volumetric images from 285 studies, with ground truth provided by the organizers. We randomly selected 20 studies for validation and 50 studies for testing, and used the remaining for training. For each study, there were four scans of T1w, T1wce, T2w and FLAIR images, and they had been co-registered”, & Page 7 – Section 4.2.1, “W-Net is a 2.5D network, and we compared using W-Net only in axial view and a fusion of axial, sagittal and coronal views. These two implementations are referred to as W- Net(A) and W-Net(ASC) respectively. The transformation parameter β in the proposed augmentation framework consisted of f l , r, s and e , where f l is a random variable for flipping along each 3D axis, r is the rotation angle along each 3D axis, s is a scaling factor and e is intensity noise. The prior distributions were: f l ∼Bern (0.5), r ∼U (0, 2 π), s ∼U (0.8, 1.2) and e ∼N (0, 0.05) according to the reduced standard deviation of a median-filtered version of a normalized image. We used this formulated augmentation during training, and also employed it to obtain TTA-based results at test time”, thus wherein the plurality of data items include generated images that are generated on the basis of different partial images included in a three-dimensional image is disclosed, because Wang teaches using volumetric images, processing axial, sagittal, and coronal views, and generating augmented/transformed images using flipping, rotation, scaling, and intensity-noise transformations. Wang’s volumetric images correspond to the three-dimensional image, Wang’s axial, sagittal, and coronal views correspond to the different partial images, and Wang’s augmented/transformed images correspond to the generated images generated on the basis of those partial images) Regarding Claim 14, Wang combined with Fukayama teaches all the limitations of claim 10 as cited above and Wang further teaches: wherein […estimated value…] is a value that indicates a position of a specific object (Wang, Page 5 – Section 4.1.1, “Two radiologists manually segmented the brain region for all the stacks slice-by-slice, where one radiologist gave a segmentation first, and then the second senior radiologist refined the segmentation if dis- agreement existed, the output of which were used as the ground truth”, & Page 5 – Section 4.1.1, “Second, the position and orientation of fetal brain have large variations, which is suitable for investigating the effect of data augmentation. For preprocessing, we normalized each stack by its intensity mean and standard deviation, and resampled each slice with pixel size 1.0 mm”, Page 7 – Section 4.2.1, “As a first demonstration of uncertainty estimation for deep learning-based brain tumor segmentation, we investigate segmentation of the whole tumor from these multi-modal images”, thus wherein […estimated value…] is a value that indicates a position of a specific object is disclosed, because Wang teaches segmentation of the fetal brain region and whole tumor from medical images. Wang’s fetal brain/whole tumor corresponds to the specific object, and Wang’s segmentation output corresponds to the estimated value because it identifies the location/position of the object in the image) Regarding Claim 15, Wang combined with Fukayama teaches all the limitations of claim 11 as cited above and Wang further teaches: wherein the estimated value is a value that indicates a position of the partial image in the three-dimensional image (Wang, Page 7 – Section 4.2.1, “We used the BraTS 2017 3 [44] training dataset that consisted of volumetric images from 285 studies, with ground truth provided by the organizers”, & Page 7 – Section 4.2.1, “W-Net is a 2.5D network, and we compared using W-Net only in axial view and a fusion of axial, sagittal and coronal views”, thus Wang teaches volumetric images and partial image views, including axial, sagittal, and coronal views, where each view corresponds to a position/orientation within the three-dimensional image. Wang’s volumetric images correspond to the three dimensional image, and Wang’s axial, sagittal, and coronal views correspond to the partial images having positions in the three-dimensional image) Regarding Claim 16, Wang teaches: receiving an input of a plurality of data items (Wang, Page 2 – Section 2.3, “The transformations for augmentation typically include flipping, cropping, rotating, and scaling training images. Abdulkadir et al. [6] and Ronneberger et al. [31] also used elastic deformations for biomedical image segmentation. Several studies have empirically found that combining predictions of multiple transformed versions of a test image helps to improve the performance”, & Page 5 – Section 4.1.1, “We collected clinical T2-weighted MRI scans of 60 fetuses in the second trimester with SSFSE on a 1.5 Tesla MR system (Aera, Siemens, Erlangen, Germany). The data for each fetus contained three stacks of 2D slices acquired in axial, sagittal and coronal views respectively, with pixel size 0.63–1.58 mm and slice thick- ness 3–6 mm. The gestational age ranged from 19 weeks to 33 weeks. We used 2640 slices from 120 stacks of 40 patients for training, 278 slices from 12 stacks of 4 patients for validation and 1180 slices from 48 stacks of 16 patients for testing”, thus receiving an input of a plurality of data items is disclosed, because Wang teaches using multiple transformed versions of a test image and multiple 2D MRI slices. The transformed test images and MRI slices correspond to the plurality of data items, and Wang’s process of using those images for prediction corresponds to receiving an input of the plurality of data items), inputting the plurality of data items to a single regression model […] from the plurality of data items (Wang, Page 2-3 – Section 2.3, “Several studies have empirically found that combining predictions of multiple transformed versions of a test image helps to improve the performance. For example, Matsunaga et al. [17] geometrically transformed test images for skin lesion classification. [32] used a single model to predict multiple transformed copies of unlabeled images for data distillation. Jin et al. [18] tested on samples extended by rotation and translation for pulmonary nodule detection”, & Page 3 – Section 3.2, “In the context of deep learning, let f ( ·) be the function represented by a neural network, and θ represent the parameters learned from a set of training images with their corresponding annotations”, thus to input the plurality of data items to a single regression model […] from the plurality of data items is disclosed, because Wang teaches using multiple transformed versions/copies of a test image and using a single model / neural network function f(·) to make predictions. The multiple transformed test images correspond to the plurality of data items, and Wang’s single model/neural network corresponds to the single regression model that receives those data items and generates prediction results from them), and integrating estimation results of […] on the basis of […] estimated by the regression model (Wang, Page 3 – Section 3.2, “Then we obtain one possible hidden image with βn and e n based on Eq. (2) , and feed it into the trained network to get its prediction, which is transformed with βn to obtain y n according to Eq. (4) . With the set Y = { y 1 , y 2 , ..., y N } sampled from p ( Y | X ), E ( Y | X ) is estimated as the average of Yand we use it as the final prediction ˆ Y for X : ˆ Y = E(Y | X ) ≈1 N N n =1 y n (8) For classification or segmentation problems, p ( Y | X ) is a discretized distribution. We obtain the final prediction for X by maximum likelihood estimation: ˆ Y = arg max y p(y | X ) ≈Mode Y”, thus and to integrate estimation results of […] on the basis of […] estimated by the regression model is disclosed, because Wang teaches feeding each hidden/transformed image into the trained network to obtain predictions y₁, y₂, …, yN, and then estimating the final prediction as the average of Y. Wang’s predictions y₁, y₂, …, yN correspond to the estimation results, and Wang’s averaging of those predictions corresponds to integrating the estimation results based on the values estimated by the regression model) Wang does not explicitly disclose […] to estimate a plurality of sets of estimated values and certainties of the estimated values […] and […] the plurality of sets […] the plurality of sets of the estimated values and the certainties of the estimated values […]. However, Fukayama teaches: […] to estimate a plurality of sets of estimated values and certainties of the estimated values […] (Fukayama, Abstract, “The estimating section calculates weights to be added the estimation results output from the regression models, based on the degrees of confidence with respect to the inputs into the regression models.”, & Background - Par. [0012], “The plurality of regression models are each capable of obtaining a probability distribution of the estimation results and a degree of confidence”, thus […] to estimate a plurality of sets of estimated values and certainties of the estimated values […] is disclosed, because Fukayama teaches regression models that output estimation results and degrees of confidence. Fukayama’s estimation results correspond to the estimated values, and Fukayama’s degrees of confidence/probability distributions correspond to the certainties of the estimated values) […] the plurality of sets […] the plurality of sets of the estimated values and the certainties of the estimated values […] (Fukayama, Description - Par. [0022], “The estimating section 4 receives an observation signal to be analyzed as an input, and obtains a mean of the estimation results (target values) and a variance in probability distributions using the regression models 21 to 2 n. Then, weights to be added to the estimation values are calculated based on the degrees of confidence (the degree of confidence is an inverse number of the variance) that are obtained by the regression models 21 to 2 n. Aggregation of weighted estimation results is performed by summing up the weighted estimated results or calculating a weighted sum”, & Description - Par. [0088], “Finally, the estimated values obtained from the individual regression models were aggregated, based on the degrees of confidence, to obtain the VA values. The estimated values were normalized such that the sum of the values should be one (1) to obtain a weight that is proportional to an inverse number of the variance. The weighted sum of the weighted estimated values respectively obtained from the regression models was calculated. The value of the weight sum was determined as the estimation result”, thus […] the plurality of sets […] the plurality of sets of the estimated values and the certainties of the estimated values […] is disclosed, because Fukayama teaches obtaining estimated values and variances, calculating degrees of confidence from the variances, and aggregating the weighted estimation results. Fukayama’s estimated values/target values correspond to the estimated values, Fukayama’s degrees of confidence/inverse variance correspond to the certainties of the estimated values, and Fukayama’s weighted sum of the estimated results corresponds to integrating the plurality of sets based on those values and certainties) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Wang with Fukayama by modifying Wang’s single model test-time augmentation prediction process to use Fukayama’s confidence based regression estimation and aggregation. Wang teaches receiving multiple image data items, inputting transformed image copies/MRI slices to a single neural network model, obtaining multiple predictions y₁, y₂, …, yN, and integrating those predictions by averaging. Fukayama teaches a computer implemented regression estimation system having processors and memories, and further teaches obtaining estimated values with degrees of confidence and aggregating the estimated values based on those confidence values. Therefore, a POSITA would have been motivated to apply Fukayama’s confidence based weighting/aggregation to Wang’s multiple prediction results so that Wang’s final prediction would account for the reliability of each estimated value instead of treating all predictions equally, thereby improving estimation performance (Fukayama, Background - Par. [0018], “The kind of an observation signal is arbitrary. In the music emotion recognition, the observation signal is a music audio signal, and the target value for the observation signal is a music emotion value. It has been confirmed that estimation performance can be improved more than ever by calculating the estimation results (mean values) and degrees of confidence using the regression models in respect of the input music audio signal and aggregating the estimation results through the maximum likelihood estimation using the calculated estimation results and degrees of confidence”) Regarding Claim 17, Wang combined with Fukayama teaches all the limitations of claim 16 as cited above and Fukayama further teaches: a non-transitory, computer-readable tangible recording medium on which a program for causing, when read by a computer, a processor of the computer to execute the regression estimation method according to claim 16 is recorded (Fukayama, Background – Par. [0030], “In a further aspect of the present invention, there is provided a computer program recorded in a computer-readable non-transitory recording medium when the method of the present invention is implemented on a computer”, & Description – Par. [0029], “FIG. 2 illustrates a method or a program algorithm for implementing the first embodiment of FIG. 1 using a computer. The program is recorded in a computer-readable, non-transitory recording medium”, thus a non-transitory, computer-readable tangible recording medium on which a program for causing, when read by a computer, a processor of the computer to execute the regression estimation method according to claim 16 is recorded is disclosed) Claims 4, 6-7, and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (hereafter Wang, a non-patent literature reference titled “Aleatoric uncertainty estimation with test-time augmentation for medical image segmentation with convolutional neural networks”) in view of Fukayama et al. (hereinafter Fukayama) (US 10614830), and in further view of Kendall et al. (hereafter Kendall, a non-patent literature reference titled “What uncertainties do we need in bayesian deep learning for computer vision”) Regarding Claim 4, Wang combined with Fukayama teaches all the limitations of claim 2 as cited above: Wang combined with Fukayama does not explicitly teach perform variable conversion to convert [… an estimated value…] output from [… a regression model…] into a first parameter of a probability distribution model and perform variable conversion to convert a value indicating [… a certainty…] output from […regression model…] into a second parameter of the probability distribution model. However, Kendall teaches: perform variable conversion to convert [… an estimated value…] output from [… a regression model…] into a first parameter of a probability distribution model (Kendall, Page 3 – Section 2.1, “For regression tasks we often define our likelihood as a Gaussian with mean given by the model output: p(y|fW(x)) = N(fW(x),σ2), with an observation noise scalar σ”, & Page 5 – Section 3.1, “we draw model weights from the approximate posterior W ∼ q(W) to obtain a model output, this time composed of both predictive mean as well as predictive variance: [ˆy, ˆ σ2] = fW(x). where f is a Bayesian convolutional neural network parametrised by model weights W. We can use a single network to transform the input x, with its head split to predict both ˆy as well as ˆσ2”, thus perform variable conversion to convert [… an estimated value…] output from [… a regression model…] into a first parameter of a probability distribution model is disclosed, because Kendall teaches defining a Gaussian likelihood where the mean is given by the model output. Kendall’s Bayesian convolutional neural network corresponds to the regression model, Kendall’s predictive mean/model output ŷ or fW(x) corresponds to the estimated value, and using that output as the mean of the Gaussian distribution corresponds to converting the estimated value into a first parameter of a probability distribution model) perform variable conversion to convert a value indicating [… a certainty…] output from […regression model…] into a second parameter of the probability distribution model (Kendall, Page 5 – Section 3.1, “we draw model weights from the approximate posterior W ∼ q(W) to obtain a model output, this time composed of both predictive mean as well as predictive variance: [ˆy, ˆ σ2] = fW(x). where f is a Bayesian convolutional neural network parametrised by model weights W. We can use a single network to transform the input x, with its head split to predict both ˆy as well as ˆσ2”, & Page 5 – Section 3.1, “In practice, we train the network to predict the log variance, si := log ˆσ2 i: LBNN(θ) = 1 D i 1 2 exp(−si)||yi − ˆyi||2 + 1 2si. (8) This is because it is more numerically stable than regressing the variance, σ2, as the loss avoids a potential division by zero. The exponential mapping also allows us to regress unconstrained scalar values, where exp(−si) is resolved to the positive domain giving valid values for variance”, thus perform variable conversion to convert a value indicating [… a certainty…] output from […regression model…] into a second parameter of the probability distribution model is disclosed, because Kendall teaches a Bayesian convolutional neural network that outputs both a predictive mean and a predictive variance, and further teaches converting a predicted log variance into a valid variance using exponential mapping. Kendall’s Bayesian convolutional neural network corresponds to the regression model, Kendall’s log variance/predictive variance corresponds to the value indicating [… a certainty…], and Kendall’s resulting variance σ ^ 2 corresponds to the second parameter of the probability distribution model) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Wang and Fukayama with Kendall by modifying the regression estimation process of Wang/Fukayama to use Kendall’s neural-network output structure in which the model outputs both a predictive mean and a predictive variance. Wang teaches using a single neural network/regression model to process multiple image inputs and generate estimation results, and Fukayama teaches forming probability distributions and aggregating estimation results based on estimated values and confidence/variance. Kendall further teaches converting the model output into parameters of a Gaussian probability distribution by using the predictive mean as a first parameter and the predictive variance/log variance as a second parameter. Therefore, a POSITA would have been motivated to incorporate Kendall’s predictive mean and variance output structure into the Wang/Fukayama regression estimation system so that the regression model could directly produce both an estimated value and an uncertainty value usable as probability-distribution parameters, thereby improving robustness to noisy or erroneous data (Kendall, Page 5 – Section 3.2, “We observe that allowing the network to predict uncertainty, allows it effectively to temper the residual loss by exp(−si), which depends on the data. This acts similarly to an intelligent robust regression function. It allows the network to adapt the residual’s weighting, and even allows the network to learn to attenuate the effect from erroneous labels. This makes the model more robust to noisy data: inputs for which the model learned to predict high uncertainty will have a smaller effect on the loss”) Regarding Claim 6, Wang and Fukayama combined with Kendall teaches all the limitations of claim 4 as cited above and Fukayama further teaches: wherein the probability distribution model is a Gaussian distribution (Fukayama, Description – Par. [0019], “The distribution of estimation results can be represented as a Gaussian distribution and an inverse number of a variance in probability distributions can be interpreted as the degree of confidence for the estimation results”, & Description – Par. [0060], “N probability distributions are obtained for the estimated value y as follows”, and Expression (9), “Pn(y)=N(yn, σn²), n=1, . . . , N”, thus wherein the probability distribution model is a Gaussian distribution is disclosed, because Fukayama teaches representing the distribution of estimation results as a Gaussian distribution. Fukayama’s Gaussian distribution Pn(y)=N(yn, σn²) corresponds to the probability distribution model) Regarding Claim 7, Wang combined with Fukayama teaches all the limitations of claim 2 as cited above and Fukayama further teaches: […] corresponding to the probability distributions of the plurality of sets during the integration (Fukayama, Description – Par. [0060], “N probability distributions are obtained for the estimated value y as follows”, and Expression (9), “ P n ( y ) = N ( y n , σ n 2 ) , n = 1 , … , N ”, & Fukayama, Description – Par. [0061], “the joint probability P J ( y ) … is obtained by calculating a product of the respective probabilities for n where n=1 to N”, thus […] corresponding to the probability distributions of the plurality of sets during the integration is disclosed, because Fukayama teaches obtaining N probability distributions P n ( y ) for the estimated value and integrating them by calculating a joint probability P J ( y ) from the product of the respective probabilities. Fukayama’s P n ( y ) = N ( y n , σ n 2 ) corresponds to the probability distributions of the plurality of sets, and Fukayama’s calculation of the joint probability P J ( y ) corresponds to the integration) and calculate a value at which […] is maximized (Fukayama, Description – Par. [0061], “the joint probability P J ( y ) … is obtained by calculating a product of the respective probabilities for n where n=1 to N”, & Fukayama, Description – Par. [0063], “A value y which maximizes the joint probability is a value of maximum likelihood estimation with respect to y ”, thus calculate a value at which […] is maximized is disclosed, because Fukayama teaches calculating a joint probability P J ( y ) from a product of respective probabilities and selecting the value y that maximizes the joint probability. Fukayama’s selected value y corresponds to the value, and Fukayama’s joint probability P J ( y ) corresponds to the probability density that is maximized) Wang combined with Fukayama does not explicitly teach perform logarithmic conversion to take a logarithm of the probability distribution, calculate a sum of logarithmic probability densities […], and […] a simultaneous logarithmic probability density […]. However, Kendall teaches: perform logarithmic conversion to take a logarithm of the probability distribution (Kendall, Page 3 – Section 2.1, “The minimisation objective is given by L ( θ , p ) = - 1 N ∑ i = 1 N l o g ⁡ p ( y i ∣ f W c i ( x i ) ) + 1 - p 2 N ∥ θ ∥ 2 ”, & Page 3 – Section 2.1, “In regression, for example, the negative log likelihood can be further simplified as - l o g ⁡ p ( y i ∣ f W c i ( x i ) ) … for a Gaussian likelihood”, thus perform logarithmic conversion to take a logarithm of the probability distribution is disclosed, because Kendall teaches taking the logarithm of the likelihood/probability distribution p ( y i ∣ f W c i ( x i ) ) . Kendall’s l o g ⁡ p ( y i ∣ f W c i ( x i ) ) corresponds to the logarithmic conversion, and Kendall’s Gaussian likelihood/probability distribution corresponds to the probability distribution) calculate a sum of logarithmic probability densities […] (Kendall, Page 3 – Section 2.1, “The minimisation objective is given by L ( θ , p ) = - 1 N ∑ i = 1 N l o g ⁡ p ( y i ∣ f W c i ( x i ) ) + 1 - p 2 N ∥ θ ∥ 2 ”, thus Kendall teaches calculating a sum of logarithmic probability densities because Kendall takes the logarithm of probability densities and sums them in the objective function) […] a simultaneous logarithmic probability density […] (Kendall, Page 3 – Section 2.1, “The minimisation objective is given by L ( θ , p ) = - 1 N ∑ i = 1 N l o g ⁡ p ( y i ∣ f W c i ( x i ) ) + 1 - p 2 N ∥ θ ∥ 2 ”, thus Kendall teaches logarithmic probability density because Kendall takes the logarithm of the likelihood/probability density) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Wang and Fukayama with Kendall by applying Kendall’s logarithmic probability-density calculation to the probability-distribution integration process of Wang/Fukayama. Wang teaches generating multiple estimation results using a trained regression/neural-network model, and Fukayama teaches obtaining N probability distributions P n ( y ) , integrating the distributions by calculating a joint probability P J ( y ) from the product of respective probabilities, and selecting the value y   that maximizes the joint probability. Kendall further teaches taking logarithms of likelihood/probability densities and calculating a summed log-likelihood. Therefore, a POSITA would have been motivated to apply Kendall’s logarithmic conversion to Fukayama’s product based joint probability calculation so that the integration could be performed as a sum of logarithmic probability densities. (Kendall, Page 5 – Section 3.2, “We observe that allowing the network to predict uncertainty, allows it effectively to temper the residual loss by exp(−si), which depends on the data. This acts similarly to an intelligent robust regression function. It allows the network to adapt the residual’s weighting, and even allows the network to learn to attenuate the effect from erroneous labels. This makes the model more robust to noisy data: inputs for which the model learned to predict high uncertainty will have a smaller effect on the loss”) Regarding Claim 9, Wang combined with Fukayama teaches all the limitations of claim 1 as cited above: Wang combined with Fukayama does not explicitly teach wherein […] is configured using a convolutional neural network. However, Kendall teaches: wherein […] is configured using a convolutional neural network (Kendall, Page 5 – Section 3.1, “ [ y ^ , σ ^ 2 ] = f W ( x ) , where f is a Bayesian convolutional neural network parametrised by model weights W ”, & Page 5 – Section 3.1, “We can use a single network to transform the input x , with its head split to predict both y ^ as well as σ ^ 2 ”, thus wherein […] is configured using a convolutional neural network is disclosed, because Kendall teaches that the model function f is a Bayesian convolutional neural network. Kendall’s Bayesian convolutional neural network corresponds to the structure configured using a convolutional neural network) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Wang and Fukayama with Kendall by configuring the regression model using Kendall’s convolutional neural network. Wang teaches processing image data using a trained neural-network model, and Fukayama teaches estimating values and confidence/probability distributions using regression models. Kendall further teaches using a Bayesian convolutional neural network f to transform input image data and output predictive values, including a predictive mean and predictive variance. Therefore, a POSITA would have been motivated to use Kendall’s convolutional neural network structure in the Wang/Fukayama regression estimation system because CNNs were known and suitable for image-based prediction tasks, and Kendall expressly applies the CNN to transform input data into regression outputs (Kendall, Page 5 – Section 3.2, “We observe that allowing the network to predict uncertainty, allows it effectively to temper the residual loss by exp(−si), which depends on the data. This acts similarly to an intelligent robust regression function. It allows the network to adapt the residual’s weighting, and even allows the network to learn to attenuate the effect from erroneous labels. This makes the model more robust to noisy data: inputs for which the model learned to predict high uncertainty will have a smaller effect on the loss”) Claims 5 is rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (hereafter Wang, a non-patent literature reference titled “Aleatoric uncertainty estimation with test-time augmentation for medical image segmentation with convolutional neural networks”) in view of Fukayama et al. (hereinafter Fukayama) (US 10614830), in view of Kendall et al. (hereafter Kendall, a non-patent literature reference titled “What uncertainties do we need in bayesian deep learning for computer vision”), and in further view of Paredes et al. (hereafter Paredes, a non-patent literature reference titled “Compressive sensing signal reconstruction by weighted median regression estimates”) Regarding Claim 5, Wang and Fukayama combined with Kendall teaches all the limitations of claim 4 as cited above: Wang and Fukayama combined with Kendall does not explicitly teach wherein [… a probability distribution model…] is a Laplace distribution. However, Paredes teaches: wherein [… a probability distribution model…] is a Laplace distribution (Paredes, Page 3 – Section III.A, “The weighted median (WM) operator has deep roots in statistical estimation theory since it emerges as the maximum likelihood (ML) estimator of location derived from a set of independent samples obeying a Laplacian distribution”, & Page 4 – Section III.A, “each element in the observation vector follows a Laplacian distribution with a common location parameter, and a (possibly) unique variance”, thus wherein [… a probability distribution model…] is a Laplace distribution is disclosed, because Paredes teaches using a Laplacian distribution for samples in a maximum likelihood estimation framework. Paredes’s Laplacian distribution corresponds to the Laplace distribution, and the distribution with a location parameter and variance corresponds to the probability distribution model) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Wang, Fukayama, and Kendall with Paredes by using Paredes’s Laplacian/Laplace distribution model in the probability-distribution-based regression estimation process. Wang teaches generating multiple estimation results from image data using a trained model, Fukayama teaches forming and integrating probability distributions based on estimated values and confidence values, and Kendall teaches using a neural network to output both a predictive mean and a predictive variance for a probability distribution model. Paredes further teaches that a Laplacian-distributed model and LAD-based estimation provide robustness to a broad class of noise, particularly heavy-tailed noise, and that the use of an l 1 -norm data-fitting term is suitable for image denoising and image restoration. Therefore, a POSITA would have been motivated to use Paredes’s Laplacian/Laplace distribution in the Wang/Fukayama/Kendall regression estimation system so that the probability distribution model would be more robust to noisy or erroneous estimation result (Paredes, Page 2 – Section I, “Unlike-regularized LS based reconstruction algorithms, the proposed-regularized LAD algorithm offers robustness to a broad class of noise, in particular to heavy tail noise [23], being optimum under the maximum likelihood (ML) principle when the underlying contamination follows a Laplacian-distributed model. Furthermore, the use of-norm in the data-fitting term has been shown to be a suitable approach for image denoising [24], image restoration [25] and sparse signal representation”) Claims 13 is rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (hereafter Wang, a non-patent literature reference titled “Aleatoric uncertainty estimation with test-time augmentation for medical image segmentation with convolutional neural networks”) in view of Fukayama et al. (hereinafter Fukayama) (US 10614830), and in further view of Igarashi et al. (hereafter Igarashi) (US 20200334818). Regarding Claim 13, Wang combined with Fukayama teaches all the limitations of claim 10 as cited above: Wang combined with Fukayama does not explicitly teach wherein […] is an elapsed time from injection of a contrast agent. However, Igarashi teaches: wherein […estimated value…] is an elapsed time from injection of a contrast agent (Igarashi, Par. [0004], “it is classified into a time phase such as an early vascular phase, an arterial predominant phase, a portal predominant phase, or a post vascular phase based on an elapsed time from the start of the injection of the contrast agent”, & Par. [0073], “the time phase information is generally classified according to the elapsed time based on the start of injection of the contrast agent”, thus wherein the estimated value is an elapsed time from injection of a contrast agent is disclosed, because Igarashi teaches generating/classifying time phase data based on elapsed time from the start of injection of a contrast agent. Igarashi time phase data/time phase information corresponds to the estimated value, and the elapsed time from the start of contrast-agent injection corresponds to the elapsed time from injection of a contrast agent) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Wang and Fukayama with Igarashi by applying the regression/probability distribution estimation process of Wang/Fukayama to Igarashi’s contrast agent medical image data to estimate time phase information based on elapsed time from injection of a contrast agent. Wang teaches processing medical image data using a trained model, and Fukayama teaches estimating values and integrating probability distributions based on estimated values and certainty/confidence values. Igarashi further teaches acquiring contrast image data and classifying the contrast image data into multiple time phases after injection of a contrast agent according to the degree of contrast of a tumor. Therefore, a POSITA would have been motivated to apply Wang/Fukayama’s estimation process to Igarashi’s contrast image data so that elapsed-time-based time phase information could be estimated from contrast medical images, thereby supporting real-time diagnosis of a tumor (Igarashi, Par. [0004], “in a contrast examination using the ultrasonic diagnostic apparatus, after injection of a contrast agent, it is classified into multiple time phases according to the degree of contrast of a tumor with the contrast agent”, & Par. [0004], “definite diagnosis of the tumor is performed in real time based on the contrast image data acquired in any of these time phases, or the multiple contrast image data acquired in the respective multiple different time phases”) Claims 18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (hereafter Wang, a non-patent literature reference titled “Aleatoric uncertainty estimation with test-time augmentation for medical image segmentation with convolutional neural networks”), and in view of Kendall et al. (hereafter Kendall, a non-patent literature reference titled “What uncertainties do we need in bayesian deep learning for computer vision”). Regarding Claim 18, Wang teaches: using training data in which data for input and a teaching signal are associated with each other, inputting the data for input to a learning model (Wang, Page 3 – Section 3.2, “In the context of deep learning, let f ( ⋅ ) be the function represented by a neural network, and θ represent the parameters learned from a set of training images with their corresponding annotations”, thus using training data in which data for input and a teaching signal are associated with each other, inputting the data for input to a learning model is disclosed, because Wang teaches learning neural-network parameters from training images and their corresponding annotations. Wang’s training images correspond to the data for input, Wang’s corresponding annotations correspond to the teaching signal, and Wang’s neural network function f ( ⋅ ) corresponds to the learning model that receives the input data) Wang does not explicitly teach obtaining an output of the estimated value and a value indicating the certainty of the estimated value from […], perform variable conversion to convert the estimated value output from […] into a first parameter of a probability distribution model, and perform variable conversion to convert the value indicating the certainty output from […] into a second parameter of the probability distribution model, calculating a loss function using the first parameter, the second parameter, and the teaching signal, and updating parameters of the learning model on the basis of a calculation result of the loss function. However, Kendall teaches: and obtaining an output of the estimated value and a value indicating the certainty of the estimated value from […a learning model…] (Kendall, Page 5 – Section 3.1, “we draw model weights from the approximate posterior W ∼ q ( W ) to obtain a model output, this time composed of both predictive mean as well as predictive variance: [ y ^ , σ ^ 2 ] = f W ( x ) ”, & Page 5 – Section 3.1, “where f is a Bayesian convolutional neural network parametrised by model weights W . We can use a single network to transform the input x , with its head split to predict both y ^ as well as σ ^ 2 ”, thus and obtaining an output of the estimated value and a value indicating the certainty of the estimated value from the learning model is disclosed, because Kendall teaches a learning model/neural network that outputs both a predictive mean and a predictive variance. Kendall’s Bayesian convolutional neural network f W ( x ) corresponds to the learning model, Kendall’s predictive mean y ^ corresponds to the estimated value, and Kendall’s predictive variance σ ^ 2 corresponds to the value indicating the certainty of the estimated value) perform variable conversion to convert the estimated value output from [… a learning model…] into a first parameter of a probability distribution model (Kendall, Page 3 – Section 2.1, “For regression tasks we often define our likelihood as a Gaussian with mean given by the model output: p(y|fW(x)) = N(fW(x),σ2), with an observation noise scalar σ”, & Page 5 – Section 3.1, “we draw model weights from the approximate posterior W ∼ q(W) to obtain a model output, this time composed of both predictive mean as well as predictive variance: [ˆy, ˆ σ2] = fW(x). where f is a Bayesian convolutional neural network parametrised by model weights W. We can use a single network to transform the input x, with its head split to predict both ˆy as well as ˆσ2”, thus perform variable conversion to convert the estimated value output from [… a learning model…] into a first parameter of a probability distribution model is disclosed, because Kendall teaches defining a Gaussian likelihood where the mean is given by the model output. Kendall’s Bayesian convolutional neural network corresponds to the regression model, Kendall’s predictive mean/model output ŷ or fW(x) corresponds to the estimated value, and using that output as the mean of the Gaussian distribution corresponds to converting the estimated value into a first parameter of a probability distribution model) perform variable conversion to convert the value indicating the certainty output from […learning model…] into a second parameter of the probability distribution model (Kendall, Page 5 – Section 3.1, “we draw model weights from the approximate posterior W ∼ q(W) to obtain a model output, this time composed of both predictive mean as well as predictive variance: [ˆy, ˆ σ2] = fW(x). where f is a Bayesian convolutional neural network parametrised by model weights W. We can use a single network to transform the input x, with its head split to predict both ˆy as well as ˆσ2”, & Page 5 – Section 3.1, “In practice, we train the network to predict the log variance, si := log ˆσ2 i: LBNN(θ) = 1 D i 1 2 exp(−si)||yi − ˆyi||2 + 1 2si. (8) This is because it is more numerically stable than regressing the variance, σ2, as the loss avoids a potential division by zero. The exponential mapping also allows us to regress unconstrained scalar values, where exp(−si) is resolved to the positive domain giving valid values for variance”, thus perform variable conversion to convert a value indicating the certainty output from […learning model…] into a second parameter of the probability distribution model is disclosed, because Kendall teaches a Bayesian convolutional neural network that outputs both a predictive mean and a predictive variance, and further teaches converting a predicted log variance into a valid variance using exponential mapping. Kendall’s Bayesian convolutional neural network corresponds to the regression model, Kendall’s log variance/predictive variance corresponds to the value indicating the certainity, and Kendall’s resulting variance σ ^ 2 corresponds to the second parameter of the probability distribution model) calculating a loss function using the first parameter, the second parameter, and the teaching signal (Kendall, Page 5 – Section 3.1, “In practice, we train the network to predict the log variance, s i : = l o g ⁡ σ ^ i 2 : L B N N ( θ ) = 1 D ∑ i 1 2 e x p ⁡ ( - s i ) ∥ y i - y ^ i ∥ 2 + 1 2 s i ”, thus calculating a loss function using the first parameter, the second parameter, and the teaching signal is disclosed, because Kendall teaches calculating a loss function using the predictive mean y ^ i , the predicted log variance s i , and the training/ground-truth value y i . Kendall’s predictive mean y ^ i corresponds to the first parameter, Kendall’s log variance s i /variance σ ^ i 2 corresponds to the second parameter, and Kendall’s ground-truth value y i corresponds to the teaching signal) updating parameters of the learning model on the basis of a calculation result of the loss function (Kendall, Page 5 – Section 3.1, “In practice, we train the network to predict the log variance, s i : = l o g ⁡ σ ^ i 2 : L B N N ( θ ) = 1 D ∑ i 1 2 e x p ⁡ ( - s i ) ∥ y i - y ^ i ∥ 2 + 1 2 s i ”, & Page 6 – Section 4, “For all experiments we train with RMS-Prop with a constant learning rate of 0.001 and weight decay 10 - 4 ”, thus updating parameters of the learning model on the basis of a calculation result of the loss function is disclosed, because Kendall teaches training the neural network using the loss function L B N N ( θ ) , where θ represents the model parameters, and further teaches training with RMS-Prop. Kendall’s θ corresponds to the parameters of the learning model, and Kendall’s training based on L B N N ( θ ) corresponds to updating the parameters on the basis of the calculation result of the loss function) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Wang with Kendall by modifying Wang’s neural-network learning model to output both an estimated value and an uncertainty/certainty value, as taught by Kendall. Wang teaches training a neural network using training images and corresponding annotations, and Wang is directed to uncertainty estimation for deep-learning-based medical image segmentation. Kendall further teaches a Bayesian convolutional neural network having a split output head that predicts both a predictive mean and a predictive variance, where the predictive mean corresponds to the estimated value and the predictive variance/log variance corresponds to the certainty value. Kendall also teaches converting those outputs into parameters of a Gaussian probability distribution, calculating a loss function using the predictive mean, predicted variance/log variance, and ground-truth value, and training the network using that loss function. Therefore, a POSITA would have been motivated to incorporate Kendall’s predictive mean/variance output structure and loss-based training into Wang’s neural-network training process so that Wang’s model could directly learn both an estimated output and an associated certainty/uncertainty value from the labeled training data. This would improve robustness and reliability of the learned model, especially for medical image analysis (Kendall, Page 5 – Section 3.2, “We observe that allowing the network to predict uncertainty, allows it effectively to temper the residual loss by exp(−si), which depends on the data. This acts similarly to an intelligent robust regression function. It allows the network to adapt the residual’s weighting, and even allows the network to learn to attenuate the effect from erroneous labels. This makes the model more robust to noisy data: inputs for which the model learned to predict high uncertainty will have a smaller effect on the loss”) Regarding Claim 20, Wang combined with Kendall teaches all the limitations of claim 18 as cited above and Kendall further teaches: wherein the probability distribution model is a Gaussian distribution, and in a case where the first parameter is μ , the second parameter is σ 2 , and the teaching signal is t , the following expression is used as the loss function: l o g ⁡ σ 2 + ( t - μ ) 2 / 2 σ 2 (Kendall, Page 3 – Section 2.1, “For regression tasks we often define our likelihood as a Gaussian with mean given by the model output: p ( y ∣ f W ( x ) ) = N ( f W ( x ) , σ 2 ) ”, & Page 5 – Section 3.1, “we train the network to predict the log variance, s i : = l o g ⁡ σ ^ i 2 : L B N N ( θ ) = 1 D ∑ i 1 2 e x p ⁡ ( - s i ) ∥ y i - y ^ i ∥ 2 + 1 2 s i ”, thus wherein the probability distribution model is a Gaussian distribution, and in a case where the first parameter is μ , the second parameter is σ 2 , and the teaching signal is t and l o g ⁡ σ 2 + ( t - μ ) 2 / 2 σ 2 is used as the loss function is disclosed, because Kendall teaches a Gaussian probability distribution having a mean given by the model output and a variance parameter. Kendall’s model output/predictive mean y ^ i corresponds to the first parameter μ , Kendall’s predictive variance σ ^ i 2 corresponds to the second parameter σ 2 , Kendall’s ground-truth value y i corresponds to the teaching signal t , and Kendall’s loss includes both a log variance term and a squared residual divided by variance term, corresponding to the Gaussian negative log-likelihood expression l o g ⁡ σ 2 + ( t - μ ) 2 / 2 σ 2 , up to scaling/constant terms) Claims 19 is rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (hereafter Wang, a non-patent literature reference titled “Aleatoric uncertainty estimation with test-time augmentation for medical image segmentation with convolutional neural networks”), and in view of Kendall et al. (hereafter Kendall, a non-patent literature reference titled “What uncertainties do we need in bayesian deep learning for computer vision”), and in further view of Paredes et al. (hereafter Paredes, a non-patent literature reference titled “Compressive sensing signal reconstruction by weighted median regression estimates”). Regarding Claim 19, Wang combined with Kendall teaches all the limitations of claim 18 as cited above and Kendall further teaches: in a case where the first parameter is μ, the second parameter is b, and the teaching signal is t (Kendall, Page 5 – Section 3.1, “ [ y ^ , σ ^ 2 ] = f W ( x ) ”, & Page 5 – Section 3.1, “We can use a single network to transform the input x , with its head split to predict both y ^ as well as σ ^ 2 ”, thus in a case where the first parameter is μ , the second parameter is b , and the teaching signal is t is disclosed, because Kendall teaches that the learning model outputs a predictive mean and an uncertainty/variance-related value, and uses a ground-truth training value in the loss function. Kendall’s predictive mean y ^ corresponds to the first parameter μ , Kendall’s uncertainty/variance-related value corresponds to the second parameter b , and Kendall’s ground-truth value y i corresponds to the teaching signal t ), the following expression is used as the loss function: l o g ⁡ b + ∣ t - μ ∣ / b (Kendall, Page 6 – Section 4, “we derive the loss function using a Laplacian prior, as opposed to the Gaussian prior used for the derivations in §3. This is because it results in a loss function which applies a L1 distance on the residuals. Typically, we find this to outperform L2 loss for regression tasks in vision. We model the benefit of combining both epistemic uncertainty as well as aleatoric uncertainty using our developments”, thus l o g ⁡ b + ∣ t - μ ∣ / b is used as the loss function is disclosed, because Kendall teaches deriving a regression loss using a Laplacian prior and applying an L1 distance to the residual. The residual between the teaching signal and predicted value corresponds to ∣ t - μ ∣ , and the Laplace negative log-likelihood includes the scale term log ⁡ b and normalized residual term ∣ t - μ ∣ / b , up to omitted constants) Wang combined with Kendall does not explicitly disclose wherein […] is a Laplace distribution. However, Paredes teaches: wherein [… a probability distribution model…] is a Laplace distribution (Paredes, Page 3 – Section III.A, “The weighted median (WM) operator has deep roots in statistical estimation theory since it emerges as the maximum likelihood (ML) estimator of location derived from a set of independent samples obeying a Laplacian distribution”, & Page 4 – Section III.A, “each element in the observation vector follows a Laplacian distribution with a common location parameter, and a (possibly) unique variance”, thus wherein [… a probability distribution model…] is a Laplace distribution is disclosed, because Paredes teaches using a Laplacian distribution for samples in a maximum likelihood estimation framework. Paredes’s Laplacian distribution corresponds to the Laplace distribution, and the distribution with a location parameter and variance corresponds to the probability distribution model) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Wang and Kendall with Paredes by using Paredes’s Laplacian/Laplace distribution model in the probability-distribution-based regression estimation process. Wang teaches generating multiple estimation results from image data using a trained model. Paredes further teaches that a Laplacian-distributed model and LAD-based estimation provide robustness to a broad class of noise, particularly heavy-tailed noise, and that the use of an l 1 -norm data-fitting term is suitable for image denoising and image restoration. Therefore, a POSITA would have been motivated to use Paredes’s Laplacian/Laplace distribution in the Wang/Kendall regression estimation system so that the probability distribution model would be more robust to noisy or erroneous estimation result (Paredes, Page 2 – Section I, “Unlike-regularized LS based reconstruction algorithms, the proposed-regularized LAD algorithm offers robustness to a broad class of noise, in particular to heavy tail noise [23], being optimum under the maximum likelihood (ML) principle when the underlying contamination follows a Laplacian-distributed model. Furthermore, the use of-norm in the data-fitting term has been shown to be a suitable approach for image denoising [24], image restoration [25] and sparse signal representation”) Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. KR 102219202 is pertinent because it teaches an artificial intelligence based stroke diagnosis method that processes brain medical images, including non-contrast CT images, using image preprocessing, normalization, region of interest extraction, and artificial intelligence models. The reference further teaches classifying hemorrhage/no-hemorrhage states, determining whether large vessel occlusion exists, estimating ASPECTS using extracted regions of interest, and determining whether mechanical thrombectomy may be applied based on the estimated ASPECTS. The reference also teaches using artificial intelligence/deep learning techniques, including CNN/LSTM-based models and machine-learning-based classification or regression structures, to analyze medical image data and generate diagnostic outputs. Because applicant’s disclosure similarly concerns machine-learning-based estimation using medical image data, trained models, estimated values, and certainty/diagnostic information generated from input data, the reference is relevant to the invention but is not relied upon in the rejection. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAHLIET ADMASU whose telephone number is (571)272-0034. The examiner can normally be reached Mon-Fri, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /M.T.A./Examiner, Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Feb 27, 2024
Application Filed
Jul 20, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month