Prosecution Insights
Last updated: August 18, 2026
Application No. 17/383,860

SUPERLOSS: A GENERIC LOSS FOR ROBUST CURRICULUM LEARNING

Non-Final OA §101§103§112
Filed
Jul 23, 2021
Priority
Oct 09, 2020 — EU 20306187.4
Examiner
SITIRICHE, LUIS A
Art Unit
2100
Tech Center
2100 — Computer Architecture & Software
Assignee
NAVER Corporation
OA Round
5 (Non-Final)
78%
Grant Probability
Favorable
5-6
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
368 granted / 474 resolved
+22.6% vs TC avg
Strong +21% interview lift
Without
With
+21.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 7m
Avg Prosecution
12 currently pending
Career history
496
Total Applications
across all art units

Statute-Specific Performance

§101
23.2%
-16.8% vs TC avg
§103
40.9%
+0.9% vs TC avg
§102
13.6%
-26.4% vs TC avg
§112
12.9%
-27.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 474 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments The Applicant’s arguments regarding the rejection of above claims have been fully considered. In reference to Applicant’s arguments about: Rejection under 35 USC 101. Examiner’s response: Applicant’s arguments regarding the previous rejection have been fully considered, but are not persuasive. Applicant asserts that the claim limitations (specifically independent claims 1, 19-20) are similar to the limitations at Example 39 of the USPTO 2019 PEG, and further asserts that the instant claim limitations are based on mathematical concepts but do not recite any mathematical relationships, formulas or calculations; however, Examiner respectfully disagrees. First, Examiner would like to provide clarification about the analysis for Example 39. The claim of Example 39 is found eligible because it does not recite an abstract idea, as the limitations (as written) does not recite any mathematical calculations; however, the claims in the instant application do recite the abstract ideas of Mathematical Concepts (as explained below in the updated rejection in this Office Action). Claim 1 (and analogous claims 19-20) explicitly recites the limitations “computing a first loss”, and “computing a weight value” which are directed to mathematical calculations. Based on the broadest reasonable interpretation, it is concluded that these limitations recite a mathematical concept, rather than just being based on a mathematical concept. For this reason, examiner understands that Example 39 is not equivalent, therefore, rejections under 35 USC 101 are still maintained. In reference to Applicant’s arguments about: 35 USC 112 rejections. Examiner’s response: Rejections under 35 USC 112 (b) are withdrawn in view of applicant’s arguments, however, rejections under 35 USC 112 (d) are maintained. Dependent claim 18 recites “The neural network of claim 1 trained according to the method of claim 1”, and this claim depends on claim 1 which previously recites “training the neural network with the set of labelled data samples according to their respective weight value”. The “training” of the neural network is already recited at independent Claim 1, and dependent claim 18 fails to provide any further details on how this training is executed, therefore, it fails to further limit the subject matter of the claim upon which it depends. In reference to Applicant’s arguments about: 35 USC 103 rejections. Examiner’s response: Applicant arguments have been fully considered but are moot in view of new grounds of rejections. Claim Objections Claims 13, 16-17 are objected to because of the following informalities: Claim 13 is missing a comma: “The method of claim 9 wherein automatically…” should read “The method of claim 9, wherein automatically…”. Claim 16 is missing a comma: “The method of claim 1 wherein…” should read “The method of claim 1, wherein…”. Claim 17 is missing a comma: “The method of claim 1 wherein…” should read “The method of claim 1, wherein…”. Appropriate correction is required. Applicant is advised that should claim 1 be found allowable, claim 20 will be objected to under 37 CFR 1.75 as being a substantial duplicate thereof. When two claims in an application are duplicates or else are so close in content that they both cover the same thing, despite a slight difference in wording, it is proper after allowing one claim to object to the other as being a substantial duplicate of the allowed claim. See MPEP § 608.01(m). Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(d): (d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph: Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. Claim 18 is rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. The claim recites: “The neural network of claim 1 trained according to the method of claim 1”. The training of the neural network is already recited at independent Claim 1, therefore, it does not further limit the subject matter at Independent claim 1 as claim 18 does not provide further details about the training of the neural network. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 stand rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception without significantly more. Step 1 analysis: In the instant case, the claims are directed to methods and a system. Thus, each of the claims falls within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). Step 2A analysis: Based on the claims being determined to be within of the four categories (Step 1), it must be determined if the claims are directed to a judicial exception (i.e., law of nature, natural phenomenon, and abstract idea), in this case the claims fall within the judicial exception of an abstract idea. Specifically the abstract ideas of Mathematical Concepts (including mathematical relationships, formulas, and/or calculations). Independent Claim 1 (and Claim 20 as being analogous): Step 2A: Prong 1 analysis: “for each data sample of a set of labeled data samples: by a first loss function for the data processing task, computing a first loss for that data sample”- this limitation recites computing a first loss, and computing a loss amounts to a mathematical calculation; being a mathematical concept/abstract idea; “by a second loss function, automatically computing a weight value for the data sample based on the first loss, the weight value indicative of a reliability of a label of the data sample predicted by the neural network for the data sample and dictating the extent to which that data sample impacts training of the neural network”- this limitation recites computing a weight value, and computing a weight value that reflects reliability and impact based on that value itself amounts to a mathematical calculation and mathematical relationships; being a mathematical concept/abstract idea. Step 2A: Prong 2 analysis: This judicial exception is not integrated into a practical application because it only recites these additional elements: “training the neural network with the set of labelled data samples according to their respective weight value” – this limitation recites using the output of the mathematical calculations (the computed weighted value) for neural network training, which under broadest reasonable interpretation in light of the specification, it amounts to applying the judicial exception to the field of use of training neural networks to perform a task, such as an image processing tasks (see MPEP 2106.05(h)). Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are directed to an abstract idea. Step 2B analysis: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element recited at Claims 1 and 20 above amounts to applying the judicial exception to a field of use. Dependent claims 2-18, when analyzed as a whole are held to be patent ineligible under 35 U.S.C. 101 because the additional recited limitation(s) fail(s) to establish that the claim(s) is/are not directed to an abstract idea. The claims are reciting further embellishment of the judicial exception. Claim 2: this claim recites further embellishment about the mathematical computation of the weight value, which amounts to further mathematical concepts, being abstract ideas. Claim 3: this claim recites further embellishment about the mathematical computation of the weight value, which amounts to further mathematical concepts, being abstract ideas. Claim 4: this claim recites a computation of the threshold value, which amounts to further mathematical concepts, being abstract ideas. Claim 5: this claim recites a computation of the threshold value using an exponential running average and a smoothing parameter, which amounts to further mathematical concepts, being abstract ideas. Claim 6: this claim recites a threshold value being a fixed value, which amounts to further mathematical concepts, being abstract ideas. Claim 7: this claim recites further embellishment about how the mathematical computation of the weight is accomplished using a regularization parameter and a threshold value, which amounts to further mathematical concepts, being abstract ideas. Claim 8: this claim recites a mathematical equation for computing the weight value, which amounts to further mathematical concepts, being abstract ideas. Claim 9: this claim recites further embellishment about the weight computation based on a confidence value, which amounts to further mathematical concepts, being abstract ideas. Claim 10: this claim recites computing the confidence value, which amounts to further mathematical concepts, being abstract ideas. Claim 11: this claim recites further embellishment about the computation of the confidence value, which amounts to further mathematical concepts, being abstract ideas. Claim 12: this claim recites a mathematical equation for the computation of the confidence value, which amounts to further mathematical concepts, being abstract ideas. Claim 13: this claim recites further embellishment about the computation of the weight value using a mathematical equation, which amounts to further mathematical concepts, being abstract ideas. Claim 14: this claim recites further embellishment about the computation of the weight value using a mathematical equation, which amounts to further mathematical concepts, being abstract ideas. Claim 15: this claim recites further embellishment about the computation of the weight value using a mathematical equation, which amounts to further mathematical concepts, being abstract ideas. Claim 16: this claim recites the use of a mathematical function, which amounts to further mathematical concepts, being abstract ideas. Claim 17: this claim recites the use of a mathematical function, which amounts to further mathematical concepts, being abstract ideas. Claim 18: this limitation recites using the output of the mathematical calculations (the computed weighted value) for neural network training, which under broadest reasonable interpretation in light of the specification, it amounts to applying the judicial exception to the field of use of training neural networks to perform a task, such as an image processing tasks (see MPEP 2106.05(h)). Viewed as a whole, these additional claim element(s) do not provide meaningful limitation(s) to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to significantly more than the abstract idea itself. Therefore, the claim(s) are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. Independent Claim 19: Step 2A: Prong 1 analysis: “for each data sample of a set of labeled data samples: using a first loss function for the data processing task, computing a first loss for that data sample”- this limitation recites computing a first loss, and computing a loss amounts to a mathematical calculation; being a mathematical concept/abstract idea; “using a second loss function, automatically computing a weight value for the data sample based on the first loss, the weight value indicative of a reliability of a label of the data sample predicted by the neural network for the data sample”- this limitation recites computing a weight value, and computing a weight value that reflects reliability based on that value itself amounts to a mathematical calculation and mathematical relationships; being a mathematical concept/abstract idea. Step 2A: Prong 2 analysis: This judicial exception is not integrated into a practical application because it only recites these additional elements: “one or more processors; memory including instructions that, when executed by the one or more processors”- these generic computer components are recited at a high level of generality such that it amounts no more than mere instructions to apply the judicial exception using a computer (see MPEP 2106.05(f)); “selectively updating a trainable parameter of the neural network based on the weight value” – this limitation recites using the output of the mathematical calculations (the computed weighted value) for neural network training updates, which under broadest reasonable interpretation in light of the specification, it amounts to applying the judicial exception to the field of use of training neural networks to perform a task, such as an image processing tasks (see MPEP 2106.05(h)). Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are directed to an abstract idea. Step 2B analysis: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements recited at Claim 19 above amounts to generic computer components are recited at a high level of generality such that it amounts no more than mere instructions to apply the judicial exception using a computer, and applying the judicial exception to a field of use. Independent Claim 20 is analogous to claim 1, as stated above. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Yao et al (US Pub. No. 2021/0133518- hereinafter Yao) in view of Guo et al (NPL: “Dynamic Task Prioritization for Multitask Learning”- hereinafter Guo). Referring to Claim 1, Yao teaches a computer-implemented method for training a neural network to perform a data processing task (see Yao at [0012]: “the techniques described herein may be used to train CNN based object detection systems to operate with improved accuracy”. Further, at [0014]: “For example, the Faster-R CNN may be a regional proposal network that can share full-image convolutional features with the detection network”) comprising: for each data sample of a set of labeled data samples: by a first loss function for the data processing task, computing a first loss for that data sample (see Yao at [0027]: “For example, the multi-scale hard miner can calculate the multi-task loss score for each candidate sample based on a localization score and a classification score corresponding to classification and localization losses calculated for each candidate sample in a respective Stochastic Gradient Descent (SGD)”. Therefore, this multi-task loss score for each candidate sample is interpreted as “a first loss for that data sample”); and training the neural network with the set of labelled data samples according to their respective weight value (see Yao at [0028]: “the predetermined number of selected sample candidates may be used to jointly train a region proposal network and a detection network”). However, Yao fails to teach: by a second loss function, automatically computing a weight value for the data sample based on the first loss, the weight value indicative of a reliability of a label of the data sample predicted by the neural network for the data sample and dictating the extent to which that data sample impacts training of the neural network; and training the neural network with the set of labelled data samples according to their respective weight value. Guo teaches, in an analogous system, by a second loss function, automatically computing a weight value for the data sample based on the first loss, the weight value indicative of a reliability of a label of the data sample predicted by the neural network for the data sample and dictating the extent to which that data sample impacts training of the neural network (see Guo at p. 6 at bottom: “We define the task-specific loss function as L∗ t(·)=FL(pc;γ0), where each example is weighted by its difficulty. The loss L∗ t(·) effectively scales the example level weight because difficult examples now contribute more to the overall loss. As a result they are given more “weight" during backpropagation. This is in line with our overall motivation: we wish to dynamically adjust the training procedure such that learning resources are not constantly allocated to easy examples”. Further, at p. 11: “Increasing the focusing parameter exponentially marks easier examples and tasks as unimportant. By increasing γ0, performance for classification and segmentation decrease…A larger γ0 forces the model to focus on detection and pose estimation, but unfortunately at the cost of classification and segmentation performance”. Therefore, this dynamic adjustment of weights corresponds to “computing a weight value”, as the weight reflects the performance (interpreted as reliability), and reflects the contribution to the loss (interpreted as the impact to the training)); and training the neural network with the set of labelled data samples according to their respective weight value (see Guo at p. 6 at bottom: “We define the task-specific loss function as L∗ t(·)=FL(pc;γ0), where each example is weighted by its difficulty. The loss L∗ t(·) effectively scales the example level weight because difficult examples now contribute more to the overall loss. As a result they are given more “weight" during backpropagation. This is in line with our overall motivation: we wish to dynamically adjust the training procedure such that learning resources are not constantly allocated to easy examples”. Therefore, this training adjustment based on the weight adjustment corresponds to training the neural network according to the weight values). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Yao with the above teachings of Guo by training a neural network to perform a data processing task using a first loss function to compute a first loss, as taught by Yao, and further computing weight adjustments for the data samples indicating reliability and impact to the training of the neural network, as taught by Guo. The modification would have been obvious because one of ordinary skill in the art would be motivated to perform dynamic task prioritization for multitask learning, thereby improving the training in curriculum learning (as suggested by Guo at section 2: “Our work on multitask learning is related to curriculum learning, which was proposed by Elman [29] to improve the training of multiple task subsets with a constant underlying distribution, starting with smaller and simpler tasks first”). Referring to Claim 18, the combination of Yao and Guo teaches the neural network of claim 1 trained according to the method of claim 1 (see Yao at [0028]: “the predetermined number of selected sample candidates may be used to jointly train a region proposal network and a detection network”. Further, see Guo at p. 6 at bottom: “We define the task-specific loss function as L∗ t(·)=FL(pc;γ0), where each example is weighted by its difficulty. The loss L∗ t(·) effectively scales the example level weight because difficult examples now contribute more to the overall loss. As a result they are given more “weight" during backpropagation. This is in line with our overall motivation: we wish to dynamically adjust the training procedure such that learning resources are not constantly allocated to easy examples”. Therefore, this training adjustment based on the weight adjustment corresponds to training the neural network according to the weight values). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Yao with the above teachings of Guo by training a neural network to perform a data processing task using a first loss function to compute a first loss, as taught by Yao, and further computing weight adjustments for the data samples indicating reliability and impact to the training of the neural network, as taught by Guo. The modification would have been obvious because one of ordinary skill in the art would be motivated to perform dynamic task prioritization for multitask learning, thereby improving the training in curriculum learning (as suggested by Guo at section 2: “Our work on multitask learning is related to curriculum learning, which was proposed by Elman [29] to improve the training of multiple task subsets with a constant underlying distribution, starting with smaller and simpler tasks first”). Referring to independent Claims 19 and 20, they are rejected on the same basis as independent claim 1, mutatis mutandis, since they are analogous claims. Claims 2-3 are rejected under 35 U.S.C. 103 as being unpatentable over Yao et al (US Pub. No. 2021/0133518- hereinafter Yao) in view of Guo et al (NPL: “Dynamic Task Prioritization for Multitask Learning”- hereinafter Guo) and further in view of Dognin et al (US Pub. No. 2015/0161988- hereinafter Dognin). Referring to Claim 2, the combination of Yao and Guo teaches the method of claim 1, however, it fails to teach wherein automatically computing the weight value for the data sample includes increasing the weight value for the data sample if the first loss is less than a threshold value. Dognin teaches, in an analogous system, wherein automatically computing the weight value for the data sample includes increasing the weight value for the data sample if the first loss is less than a threshold value (see Dognin at [0050]: “In accordance with an embodiment of the present invention, weights are chosen based on a difference between a held-out loss of a current iteration and a held-out loss of at least one previous iteration. The assigned weight is inversely proportional to the difference, such that a relatively larger difference between a held-out loss of a current iteration and a held-out loss of a previous iteration results in a smaller weight given to the gradient corresponding to the held-out loss of the previous iteration, and a relatively smaller difference results in a larger weight given to the gradient corresponding to the held-out loss of the previous iteration”. Further, at [0075]: “According to an embodiment, the weighting component 212 uses a weighting function to assign higher weights to gradients with held-out loss values closer to a held-out loss value of a current model, and lower weights to gradients with held-out loss values farther from the held-out loss value of the current model”. Therefore, the held-out loss value of the current model is interpreted as the threshold, as based on the difference of the loss and this value, the weights are increased or decreased, as they are inversely proportional (high loss would amount to lower the weight, and vice-versa)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Yao and Guo with the above teachings of Dognin by training a neural network to perform a data processing task using a first loss function to compute a first loss further computing weight adjustments for the data samples indicating reliability and impact to the training of the neural network, as taught by Yao and Guo, wherein the weight adjustments include increasing the weight value according to a threshold of the loss, as taught by Dognin. The modification would have been obvious because one of ordinary skill in the art would be motivated to dynamically adjust the weights according to the loss for optimizing the neural network parameters (as suggested by Dognin at [0004]: “A component of the DNN training procedure is sequence training (ST), where the network parameters are optimized under a sequence classification criterion” and [0006]: “This procedure, referred to herein as dynamic stochastic average gradient with Hessian-free (DSAG-HF), leverages gradient averaging as proposed in a stochastic average gradient approach, and carries out an HF conjugate gradient (CG)-based optimization using these averaged gradients”). Referring to Claim 3, the combination of Yao, Guo and Dognin teaches the method of claim 2, wherein automatically computing the weight value for the data sample includes decreasing the weight value for the data sample if the first loss is greater than the threshold value (see Dognin at [0050]: “In accordance with an embodiment of the present invention, weights are chosen based on a difference between a held-out loss of a current iteration and a held-out loss of at least one previous iteration. The assigned weight is inversely proportional to the difference, such that a relatively larger difference between a held-out loss of a current iteration and a held-out loss of a previous iteration results in a smaller weight given to the gradient corresponding to the held-out loss of the previous iteration, and a relatively smaller difference results in a larger weight given to the gradient corresponding to the held-out loss of the previous iteration”. Further, at [0075]: “According to an embodiment, the weighting component 212 uses a weighting function to assign higher weights to gradients with held-out loss values closer to a held-out loss value of a current model, and lower weights to gradients with held-out loss values farther from the held-out loss value of the current model”. Therefore, the held-out loss value of the current model is interpreted as the threshold, as based on the difference of the loss and this value, the weights are increased or decreased, as they are inversely proportional (high loss would amount to lower the weight, and vice-versa)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Yao and Guo with the above teachings of Dognin by training a neural network to perform a data processing task using a first loss function to compute a first loss further computing weight adjustments for the data samples indicating reliability and impact to the training of the neural network, as taught by Yao and Guo, wherein the weight adjustments include decreasing the weight value according to a threshold of the loss, as taught by Dognin. The modification would have been obvious because one of ordinary skill in the art would be motivated to dynamically adjust the weights according to the loss for optimizing the neural network parameters (as suggested by Dognin at [0004]: “A component of the DNN training procedure is sequence training (ST), where the network parameters are optimized under a sequence classification criterion” and [0006]: “This procedure, referred to herein as dynamic stochastic average gradient with Hessian-free (DSAG-HF), leverages gradient averaging as proposed in a stochastic average gradient approach, and carries out an HF conjugate gradient (CG)-based optimization using these averaged gradients”). Claims 4-5 are rejected under 35 U.S.C. 103 as being unpatentable over Yao et al (US Pub. No. 2021/0133518- hereinafter Yao) in view of Guo et al (NPL: “Dynamic Task Prioritization for Multitask Learning”- hereinafter Guo), in view of Dognin et al (US Pub. No. 2015/0161988- hereinafter Dognin), and further in view of Tam Nguyen et al (NPL: “Self: Learning To Filter Noisy Labels With Self Ensembling”- as submitted in IDS 08/23/2021, hereinafter Tam). Referring to Claim 4, the combination of Yao, Guo and Dognin teaches the method of claim 2, however, fails to teach further comprising computing the threshold value based on a running average of the first loss. Tam teaches, in an analogous system, further comprising computing the threshold value based on a running average of the first loss (see Tam at Abstract: “For the filtering, we form running averages of predictions over the entire training dataset using the network output at different training epochs. We show that these ensemble estimates yield more accurate identification of inconsistent predictions throughout training than the single estimates of the network at the most recent training epoch”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Yao, Guo and Dognin with the above teachings of Tam by training a neural network to perform a data processing task using a first loss function to compute a first loss further computing weight adjustments based on a threshold for the data samples indicating reliability and impact to the training of the neural network, as taught by Yao, Guo and Dognin, wherein the threshold is computed based on a running average of the loss, as taught by Tam. The modification would have been obvious because one of ordinary skill in the art would be motivated to yield more accurate identification of inconsistent predictions throughout training than the single estimates of the network at the most recent training epoch (as suggested by Tam at Abstract). Referring to Claim 5, the combination of Yao, Guo and Dognin teaches the method of claim 2, however, fails to teach further comprising computing the threshold value based on an exponential running average of the first loss and using a smoothing parameter. Tam teaches, in an analogous system, further comprising computing the threshold value based on an exponential running average of the first loss and using a smoothing parameter (see Tam at pp. 4-5, section 2.3: “Model ensemble with Mean Teacher A natural way to form a model ensemble is by using an exponential running average of model snapshots (Fig. 3a). This idea was proposed in Tarvainen & Valpola (2017) for semi-supervised learning and is known as the Mean Teacher model. In our framework, both the mean teacher model and the normal model are evaluated on all data to preserve the consistency between both models. The consistency loss between student and teacher output distribution can be realized with Mean-Square-Error loss or Kullback-Leibler-divergence” and “Prediction ensemble Additionally, we propose to collect the sample predictions over multiple training epochs: zj = αzj−1 + (1 − α)ˆzj , whereby zj depicts the moving-average prediction of sample k at epoch j, α is a momentum, ˆzj is the model prediction for sample k in epoch j. This scheme is displayed in Fig. 3b. For each sample, we store the moving-average predictions, accumulated over the past iterations. Besides having a more stable basis for the filtering step, our proposed procedure also leads to negligible memory and computation overhead”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Yao, Guo and Dognin with the above teachings of Tam by training a neural network to perform a data processing task using a first loss function to compute a first loss further computing weight adjustments based on a threshold for the data samples indicating reliability and impact to the training of the neural network, as taught by Yao, Guo and Dognin, wherein the threshold is computed based on a running average of the loss, as taught by Tam. The modification would have been obvious because one of ordinary skill in the art would be motivated to make the training more stable, leading to negligible memory and computation overhead (as suggested by Tam at section 2.3). Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Yao et al (US Pub. No. 2021/0133518- hereinafter Yao) in view of Guo et al (NPL: “Dynamic Task Prioritization for Multitask Learning”- hereinafter Guo), in view of Dognin et al (US Pub. No. 2015/0161988- hereinafter Dognin), and further in view of Anisimov et al (US Patent 11,468,288- hereinafter Anisimov). Referring to Claim 6, the combination of Yao, Guo and Dognin teaches the method of claim 2, however, fails to teach wherein the threshold value is a fixed predetermined value. Anisimov teaches, in an analogous system, wherein the threshold value is a fixed predetermined value (see Anisimov at Claim 16: “the neural network functions include an optimizing a loss function by of one of a mean squared error and a cross-entropy function, such that the recurrent neural network applies weights to input data, the loss function is calculated, the loss function is optimized by changing the weights within the recurrent neural network, the recurrent neural network repeats the process until a predetermined threshold is reached”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Yao, Guo and Dognin with the above teachings of Anisimov by training a neural network to perform a data processing task using a first loss function to compute a first loss further computing weight adjustments based on a threshold for the data samples indicating reliability and impact to the training of the neural network, as taught by Yao, Guo and Dognin, wherein the threshold is a predetermined threshold, as taught by Anisimov. The modification would have been obvious because one of ordinary skill in the art would be motivated to optimize the loss function by changing weights iteratively until a predetermined threshold is met (as suggested by Anisimov at Claim 16). Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Yao et al (US Pub. No. 2021/0133518- hereinafter Yao) in view of Guo et al (NPL: “Dynamic Task Prioritization for Multitask Learning”- hereinafter Guo), in view of Luketina et al (NPL: “ Scalable Gradient-Based Tuning of Continuous Regularization Hyperparameters”- hereinafter Luketina), and further in view of Dognin et al (US Pub. No. 2015/0161988- hereinafter Dognin). Referring to Claim 7, the combination of Yao and Guo teaches the method of claim 1, wherein automatically computing the weight value includes, by the second loss function, automatically computing the weight value further based on a regularization hyperparameter and a threshold value. Luketina teaches, in an analogous system wherein automatically computing the weight value includes, by the second loss function, automatically computing the weight value further based on a regularization hyperparameter (see Luketina at 2nd page, right column: “Although the proposed method could work in principle for any continuous hyperparameter, we have specifically focused on studying tuning of regularization hyperparameters”. Further at Section 2. Proposed Method: “We propose a method, T1−T2, for tuning continuous hyper parameters of a model using the gradient of the performance of the model on a separate validation set T2. In essence, we train a neural network model on a training set T1 as usual. However, for each update of the network weights and biases, i.e. the elementary parameters of the network, we tune the hyperparameters so as to make the direction of the weight update as beneficial as possible for the validation cost on a separate dataset T2”) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Yao and Guo with the above teachings of Luketina by training a neural network to perform a data processing task using a first loss function to compute a first loss further computing weight adjustments for the data samples indicating reliability and impact to the training of the neural network, as taught by Yao and Guo, wherein the weights computation is based on a regularization hyperparameter, as taught by Luketina. The modification would have been obvious because one of ordinary skill in the art would be motivated to tune the hyperparameters so as to make the direction of the weight update as beneficial as possible for the validation cost (as suggested by Luketina at [Section 2]). Dognin teaches, in an analogous system, wherein automatically computing the weight value includes, by the second loss function, automatically computing the weight value further based on a threshold value (see Dognin at [0050]: “In accordance with an embodiment of the present invention, weights are chosen based on a difference between a held-out loss of a current iteration and a held-out loss of at least one previous iteration. The assigned weight is inversely proportional to the difference, such that a relatively larger difference between a held-out loss of a current iteration and a held-out loss of a previous iteration results in a smaller weight given to the gradient corresponding to the held-out loss of the previous iteration, and a relatively smaller difference results in a larger weight given to the gradient corresponding to the held-out loss of the previous iteration”. Further, at [0075]: “According to an embodiment, the weighting component 212 uses a weighting function to assign higher weights to gradients with held-out loss values closer to a held-out loss value of a current model, and lower weights to gradients with held-out loss values farther from the held-out loss value of the current model”. Therefore, the held-out loss value of the current model is interpreted as the threshold, as based on the difference of the loss and this value, the weights are increased or decreased, as they are inversely proportional (high loss would amount to lower the weight, and vice-versa)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Yao, Guo and Luketina with the above teachings of Dognin by training a neural network to perform a data processing task using a first loss function to compute a first loss further computing weight adjustments based on a threshold for the data samples indicating reliability and impact to the training of the neural network, as taught by Yao, Guo and Luketina, wherein the weights computation is based on a threshold, as taught by Dognin. The modification would have been obvious because one of ordinary skill in the art would be motivated to dynamically adjust the weights according to the loss for optimizing the neural network parameters (as suggested by Dognin at [0004]: “A component of the DNN training procedure is sequence training (ST), where the network parameters are optimized under a sequence classification criterion” and [0006]: “This procedure, referred to herein as dynamic stochastic average gradient with Hessian-free (DSAG-HF), leverages gradient averaging as proposed in a stochastic average gradient approach, and carries out an HF conjugate gradient (CG)-based optimization using these averaged gradients”). Claims 9-11 are rejected under 35 U.S.C. 103 as being unpatentable over Yao et al (US Pub. No. 2021/0133518- hereinafter Yao) in view of Guo et al (NPL: “Dynamic Task Prioritization for Multitask Learning”- hereinafter Guo), in view of Luketina et al (NPL: “Scalable Gradient-Based Tuning of Continuous Regularization Hyperparameters”- hereinafter Luketina), in view of Dognin et al (US Pub. No. 2015/0161988- hereinafter Dognin), and further in view of Arik et al (US Pub. No. 2021/0034976- hereinafter Arik). Referring to Claim 9, the combination of Yao, Guo, Luketina and Dognin teaches the method of claim 7, however, fails to teach wherein automatically computing the weight value includes, by the second loss function, automatically computing the weight value further based on a confidence value of the data sample. Arik teaches, in an analogous system, wherein automatically computing the weight value includes, by the second loss function, automatically computing the weight value further based on a confidence value of the data sample (see Arik at [0011]: “sampling a training batch of source data samples from the source data set having a particular size; and selecting the source data samples from the training batch of source data samples that have the N-best confidence scores for use in training the deep learning model to learn the encoder weights, the source classifier layer weights, and the target classifier layer weights that minimize the loss function”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Yao, Guo, Luketina and Dognin with the above teachings of Arik by training a neural network to perform a data processing task using a first loss function to compute a first loss further computing weight adjustments based on a threshold for the data samples indicating reliability and impact to the training of the neural network, as taught by Yao, Guo and Luketina, wherein the weights computation is based on a confidence value of the data sample, as taught by Arik. The modification would have been obvious because one of ordinary skill in the art would be motivated to determine and select the samples with the best confidence score in order to minimize the loss function (as suggested by Arik at [0011]). Referring to Claim 10, the combination of Yao, Guo, Luketina, Dognin and Arik teaches the method of claim 9, further comprising computing the confidence value of the data sample based on the first loss (see Arik at [0011]: “sampling a training batch of source data samples from the source data set having a particular size; and selecting the source data samples from the training batch of source data samples that have the N-best confidence scores for use in training the deep learning model to learn the encoder weights, the source classifier layer weights, and the target classifier layer weights that minimize the loss function”. Further at [0029]: “Here, the encoding network 152 encodes input features (e.g., images) from the source dataset 104 and the source classifier layer 154 (also referred to as “source decision layer”) uses the encoded input features to output a confidence score, whereby the training objective determines a source dataset classification loss”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Yao, Guo, Luketina and Dognin with the above teachings of Arik by training a neural network to perform a data processing task using a first loss function to compute a first loss further computing weight adjustments based on a threshold for the data samples indicating reliability and impact to the training of the neural network, as taught by Yao, Guo and Luketina, wherein the weights computation is based on a confidence value of the data sample, as taught by Arik. The modification would have been obvious because one of ordinary skill in the art would be motivated to determine and select the samples with the best confidence score in order to minimize the loss function (as suggested by Arik at [0011]). Referring to Claim 11, the combination of Yao, Guo, Luketina, Dognin and Arik teaches the method of claim 9, wherein computing the confidence value of the data sample includes computing the confidence value based on minimizing the second loss function for the first loss (see Arik at [0011]: “sampling a training batch of source data samples from the source data set having a particular size; and selecting the source data samples from the training batch of source data samples that have the N-best confidence scores for use in training the deep learning model to learn the encoder weights, the source classifier layer weights, and the target classifier layer weights that minimize the loss function”. Further at [0029]: “Here, the encoding network 152 encodes input features (e.g., images) from the source dataset 104 and the source classifier layer 154 (also referred to as “source decision layer”) uses the encoded input features to output a confidence score, whereby the training objective determines a source dataset classification loss”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Yao, Guo, Luketina and Dognin with the above teachings of Arik by training a neural network to perform a data processing task using a first loss function to compute a first loss further computing weight adjustments based on a threshold for the data samples indicating reliability and impact to the training of the neural network, as taught by Yao, Guo and Luketina, wherein the weights computation is based on a confidence value of the data sample, as taught by Arik. The modification would have been obvious because one of ordinary skill in the art would be motivated to determine and select the samples with the best confidence score in order to minimize the loss function (as suggested by Arik at [0011]). Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Yao et al (US Pub. No. 2021/0133518- hereinafter Yao) in view of Guo et al (NPL: “Dynamic Task Prioritization for Multitask Learning”- hereinafter Guo) and further in view of Daemen et al (US Patent 10,776,728 - hereinafter Daemen). Referring to Claim 16, the combination of Yao and Guo teaches the method of claim 1, however, fails to teach wherein the second loss function is a monotonically increasing concave function. Daemen teaches, in an analogous system, wherein the second loss function is a monotonically increasing concave function (see Daemen at Col. 4: 51-57: “In some examples, each importance parameter (v) may be based on a monotonically increasing concave function of the corresponding target (t) (e.g., v.sub.i:=∜t.sub.i). In examples described herein, the importance parameters have positive values. Example importance parameters (v) per criteria variable for computing target loss may be expressed and/or defined as follows”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Yao and Guo with the above teachings of Daemen by training a neural network to perform a data processing task using a first loss function to compute a first loss further computing weight adjustments with a second loss, as taught by Yao and Guo, wherein the second loss is a monotonically increasing concave function, as taught by Daemen. The modification would have been obvious because one of ordinary skill in the art would be motivated to use a monotonically increasing concave function in order to determine the importance of each data sample (as suggested by Daemen at Col. 4: 51-57). Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Yao et al (US Pub. No. 2021/0133518- hereinafter Yao) in view of Guo et al (NPL: “Dynamic Task Prioritization for Multitask Learning”- hereinafter Guo) and further in view of Patton (NPL: “Volatility forecast comparison using imperfect volatility proxies” - hereinafter Patton). Referring to Claim 17, the combination of Yao and Guo teaches the method of claim 1, however, fails to teach wherein the second loss function is a homogeneous function. Patton teaches, in an analogous system, wherein the second loss function is a homogeneous function (see Patton at p. 251, right column: “The following proposition shows that by using a homogeneous robust loss function, the ranking of any two (possibly imperfect) forecasts is invariant to a re-scaling of the data. It further provides an example where the ranking can be reversed simply with a rescaling of the data if a non-homogeneous robust loss function is used”. Further, at p. 252, left column: “With the above motivation for homogeneous loss functions, we now derive the subset of homogeneous, robust loss functions. It turns out that this subset of functions is indexed by a single parameter, which determines the both degree of homogeneity and the shape of the loss function”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Yao and Guo with the above teachings of Patton by training a neural network to perform a data processing task using a first loss function to compute a first loss further computing weight adjustments with a second loss, as taught by Yao and Guo, wherein the second loss is a homogeneous function, as taught by Patton. The modification would have been obvious because one of ordinary skill in the art would be motivated to use a robust homogeneous lost function such that the ranking of any two predictions is invariant to a re-scaling of the data (as suggested by Patton at p.251). Allowable Subject Matter For claims 8, 12-15, no art rejection is made for these claims, they are only rejected under 35 USC 101 as explained above in this office action. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to LUIS A SITIRICHE whose telephone number is (571)270-1316. The examiner can normally be reached M-F 9am-6pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /LUIS A SITIRICHE/Primary Examiner, Art Unit 2126
Read full office action

Prosecution Timeline

Show 15 earlier events
Jun 13, 2025
Notice of Allowance
Dec 09, 2025
Response after Non-Final Action
Dec 17, 2025
Non-Final Rejection (signed) — §101, §103, §112
Jan 27, 2026
Non-Final Rejection mailed — §101, §103, §112
Apr 22, 2026
Response Filed
Aug 03, 2026
Non-Final Rejection mailed — §101, §103, §112
Aug 11, 2026
Examiner Interview Summary
Aug 11, 2026
Applicant Interview (Telephonic)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705301
ANALYSIS OF CLUSTERED DATA
4y 2m to grant Granted Aug 11, 2026
Patent 12682276
METHODS AND SYSTEMS FOR PROCESSING UNSTRUCTURED AND UNLABELLED DATA
4y 9m to grant Granted Jul 14, 2026
Patent 12670232
MODEL PREDICTION CONFIDENCE UTILIZING DRIFT
4y 8m to grant Granted Jun 30, 2026
Patent 12664481
Systems and Methods for Predictive Coding Utilizing Confidence Levels
5y 2m to grant Granted Jun 23, 2026
Patent 12651040
SAMPLING FROM A SET SPINS WITH CLAMPING
4y 6m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+21.4%)
3y 7m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 474 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month