Prosecution Insights
Last updated: October 02, 2026
Application No. 17/829,848

METHOD AND SYSTEM FOR MACHINE LEARNING FROM IMBALANCED DATA WITH NOISY LABELS

Non-Final OA §103
Filed
Jun 01, 2022
Priority
Nov 28, 2021 — provisional 63/283,492
Examiner
MEYER, JACQUELINE CHRISTINE
Art Unit
2144
Tech Center
2100 — Computer Architecture & Software
Assignee
NAVER Corporation
OA Round
3 (Non-Final)
65%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 65% — above average
65%
Career Allowance Rate
15 granted / 23 resolved
+10.2% vs TC avg
Strong +62% interview lift
Without
With
+61.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
13 currently pending
Career history
42
Total Applications
across all art units

Statute-Specific Performance

§101
24.4%
-15.6% vs TC avg
§103
54.7%
+14.7% vs TC avg
§102
8.9%
-31.1% vs TC avg
§112
11.1%
-28.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 23 resolved cases

Office Action

§103
DETAILED ACTION This nonfinal office action is responsive to the Appeal Brief filed on June 12, 2026. The finality of the office action dated November 26, 2025 is withdrawn. Claims 1-21 are pending. Claims 1, 18, and 20 are independent. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: "An encoder module and projection module configured to generate matrix representations in claim 18. “a training module configured to” in claim 18. These modules are being interpreted under 35 U.S.C. 112(f), the broadest reasonable interpretation would include the definition and examples given in the specification in paragraph 0102 as an ASIC; digital, analog or mixed discrete circuit; digital, analog or mixed integrated circuit; combinational logic circuit; a gate array; processor circuit; memory circuit; or other suitable hardware components. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 5-7, 11-13, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (US20210374553), hereinafter Li, in view of Chen et al. (US20210319266), hereinafter Chen, in view of Menon et al. (Long-Tail Learning via Logit Adjustment), hereinafter Menon. Menon was cited in the IDS provided on June 1, 2022. Regarding claim 1, Li teaches: A computer-implemented method for training an artificial neural network with training data including samples and corresponding labels for performing a task, the method comprising: (Li, paragraph 0024: “In some embodiments, an iterative label correction may be implemented. For example, for each sample, information from the neighbors of the respective sample may be aggregated to create a pseudo-label. A subset of training samples with confident pseudo-labels are then selected to compute the supervised losses.” – This is a method for training neural networks with training data that includes samples and labels.) …wherein the artificial neural network includes an encoder module and a projection module configured to generate the matrix representations based on ones of the samples, respectively; and (Li, paragraph 0033: “The neural network 130 includes at least three components: (1) a deep encoder, e.g., a convolutional neural network (CNN) 245 a-c (or collectively referred to as 245) that encodes an image xi or an augmented sample of image xi to a high-dimensional feature vi; (2) a classifier 255 (e.g., a fully-connected layer followed by softmax) that receives vi as input and outputs class predictions; (3) a linear autoencoder 260 a-c (or collectively referred to as 260) that projects vi into a low-dimensional embedding z i ∈ R d .” – The deep encoder is analogous to the encoder module and the autoencoder is analogous to the projection module, while the augmented sample of image xi to a high-dimensional feature vi and the low-dimensional embedding z i are analogous to the matrix representations based on one of the samples.) Li does not explicitly teach: pre-training the artificial neural network to generate matrix representations that are invariant to a predetermined set of data augmentations applied to a sample, after the pre-training, fine-tune training the artificial neural network using a loss function, wherein fine-tuning the artificial neural network includes adjusting, based on the labels, one or more weights of the projection module while maintaining constant weights of the encoder module, and wherein the loss function is based on a logit adjustment loss that is based on logits that are adjusted based on a class distribution of the training data. However, Chen teaches: pre-training the artificial neural network to generate matrix representations that are invariant to a predetermined set of data augmentations applied to a sample, (Chen, paragraph 0084: “In particular, z = g ( h ) is trained to be invariant to data transformations.” – z is analogous to the matrix representations and data transformations is analogous to data augmentations.) after the pre-training, fine-tune training the artificial neural network using a loss function, wherein fine-tuning the artificial neural network includes adjusting, based on the labels, one or more weights of the projection module while maintaining constant weights of the encoder module, and (Chen, paragraph 0006: “evaluating a loss function that evaluates a difference between the first projected representation and the second projected representation, modifying one or more values of one or more parameters of one or both of the base encoder neural network and the projection head neural network based at least in part on the loss function… performing fine-tuning of the image classification model based on a set of labeled images..” – modifying the parameters of one or both of the base encoder neural network and the projection head neural network indicates that it is capable of adjusting the weights of the projection module while maintaining constant weights to the encoder module while fine-tuning based on a set of labeled images is fine-tuning the artificial neural network based on the labels.) Chen is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and self-supervised contrastive learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li, which already teaches a neural network with an encoder and a projection module configured to generate matrix representations but does not explicitly teach that the matrix representations are invariant or that the fine-tuning includes adjusting weights of only the projection module, to include the teachings of Chen which does teach that the matrix representations are invariant or that the fine-tuning includes adjusting weights of only the projection module in order to "provide improved visual representations." (Chen, abstract) Li and Chen do not explicitly teach: wherein the loss function is based on a logit adjustment loss that is based on logits that are adjusted based on a class distribution of the training data. However, Menon teaches: wherein the loss function is based on a logit adjustment loss that is based on logits that are adjusted based on a class distribution of the training data. (Menon, section 5: “We now show how to directly bake logit adjustment into the softmax cross-entropy. We show that this approach has an intuitive relation to existing loss modification techniques.” and section 5.1: “To do so, consider the following logit adjustment softmax cross-entropy loss for τ > 0 : PNG media_image1.png 84 912 media_image1.png Greyscale Given a scorer that minimizes the above, we now predict a r g m a x y ∈ L f y ( x ) as usual. Compared to the standard softmax cross-entropy (1), the above applies a label-dependent offset to each logit.” – The label-dependent offset indicates that the logit adjustment loss (equation 10) is based on a class distribution.) Menon is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and learning under class imbalance. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li and Chen, which already teaches pre-training and fine-tuning a neural network but does not explicitly teach that the loss function used is based on a logit adjustment loss based on a class distribution of the training data, to include the teachings of Menon which does teach that the loss function used is based on a logit adjustment loss based on a class distribution of the training data as it leads to superior performance. “However, unlike most previous techniques from these streams, logit adjustment is endowed with a clear statistical grounding: by construction, the optimal solution under such adjustment coincides with the Bayes-optimal solution (7) for the balanced error, i.e., it is Fisher consistent for minimising the balanced error. We shall demonstrate this translates into superior empirical performance (§6).” (Menon, section 3, paragraph 3) “However, §3 indicates that our loss in (10) has a firm statistical grounding: it ensures Fisher consistency for the balanced error.” (Menon, section 5.2, paragraph 3) Regarding claim 5, Li, Chen, and Menon teach the method of claim 1, as cited above. Li further teaches: before fine-tune training the artificial neural network, estimating the class distribution. (Li, paragraph 0086: “As training progresses, the neural network learns to separate OOD samples from in-distribution samples, and cluster samples of the same class despite their noisy labels, as illustrated by the separated cloud clusters.” – In order for the training (fine-tuning) to be able to separate the out-of-distribution samples from in-distribution samples it would necessarily have needed to estimate the class distribution prior to this step in order to determine if the samples are OOD or in-distribution. Since this is a classification neural network the labels are analogous to the classes.) Regarding claim 6, Li, Chen, and Menon teach the method of claim 1, as cited above. Li and Chen do not explicitly teach: wherein the class distribution of the labels over the samples is a long-tailed class distribution. However, Menon further teaches: wherein the class distribution of the labels over the samples is a long-tailed class distribution. (Menon, section 2, paragraph 2: “The setting of learning under class imbalance or long-tailed learning is where the distribution P ( y ) is highly skewed, so that many (rare or “tail”) labels have a very low probability of occurrence.”) Regarding claim 7, Li, Chen, and Menon teach the method of claim 1, as cited above. Li further teaches: wherein the labels are noisy. (Li, paragraph 0031: “Diagram 200 shows that an input data sample 240, similar to one of the images 110a-n with a noisy label shown in FIG. 1, is used for noise-robust contrastive learning.”) Regarding claim 11, Li, Chen, and Menon teach the method of claim 1, as cited above. Li further teaches: wherein the pre-training includes optimizing a loss between respective representations generated by the artificial neural network for a first augmented data sample and a second augmented data sample, (Li, paragraph 0029: “On one hand, for consistency learning, the “strong” embedding 136 and the “weak” embedding 135 of the same image sample are compared to generate an unsupervised consistency loss for consistency learning 140, which is minimized to increase the consistency of neural network performance.” – The weak embedding and strong embedding are representations generated from a weak augmented data sample and strong augmented data sample, respectively, which corresponds to two separate augmented data samples.) wherein the first and second augmented samples are generated by the artificial neural network by applying first and second data augmentations of the set of predetermined data augmentations, respectively, to the sample. (Li, paragraph 0028: “The neural network 130 may project the training images 110 a-n to a low-dimensional subspace, e.g., by generating a “weak” normalized embedding 135 from a weakly augmented image sample and a “strong” normalized embedding 136 from a strongly augmented image sample from the same training image sample.” – The weakly augmented image sample and strongly augmented image sample are from the same sample and therefore are analogous to a first and second data augmentation.) Regarding claim 12, Li, Chen, and Menon teach the method of claim 1, as cited above. Li further teaches: wherein the samples are image samples, and wherein the data augmentations are image transformations. (Li, paragraph 0028: “The neural network 130 may project the training images 110 a-n to a low-dimensional subspace, e.g., by generating a “weak” normalized embedding 135 from a weakly augmented image sample and a “strong” normalized embedding 136 from a strongly augmented image sample from the same training image sample.” – The weakly and strongly augmented image samples are analogous to image transformations, see fig. 2, numbers 241 and 242.) Regarding claim 13, Li, Chen, and Menon teach the method of claim 1, as cited above. Li further teaches: further comprising, by the artificial neural network, classifying an object in an image after the fine-tune training. (Li, paragraph 0003: “For example, a neural network classifier may predict a class of input information among a predetermined set of classes, e.g., identifying that a photo of a fruit orange belongs to the class “orange.” To achieve this, neural networks learn to make predictions gradually, by a process of trial and error, using a machine learning process. A given neural network model may be trained using a large number of training examples, proceeding iteratively until the neural network model begins to consistently make similar inferences from the training examples that a human may make.” – A neural network that is trained is analogous to being fine-tuned.) Regarding claim 18, Li teaches: an artificial neural network including an encoder module and a projection module configured to generate matrix representations based on input samples; (Li, paragraph 0033: “The neural network 130 includes at least three components: (1) a deep encoder, e.g., a convolutional neural network (CNN) 245 a-c (or collectively referred to as 245) that encodes an image xi or an augmented sample of image xi to a high-dimensional feature vi; (2) a classifier 255 (e.g., a fully-connected layer followed by softmax) that receives vi as input and outputs class predictions; (3) a linear autoencoder 260 a-c (or collectively referred to as 260) that projects vi into a low-dimensional embedding z i ∈ R d .” – The deep encoder is analogous to the encoder module and the autoencoder is analogous to the projection module, while the augmented sample of image xi to a high-dimensional feature vi and the low-dimensional embedding z i are analogous to the matrix representations based on one of the samples.) training data including samples and corresponding labels; and (Li, paragraph 0058: “Method 500 starts with step 502, at which a training set of data samples may be obtained, each data sample having a noisy label.” – The data samples having noisy labels is analogous to the training data including samples and corresponding labels.) a training module configured to: (Li, paragraph 0049: “In some examples, the noise-robust contrastive learning module 330 may also handle the iterative training and/or evaluation of a system or model.” – The noise robust contrastive learning module is analogous to a training module.) Li does not explicitly teach: pre-train the artificial neural network to generate matrix representations that are invariant to a predetermined set of data augmentations applied to a sample; and after the pre-training, fine-tune train the artificial neural network using a loss function, the fine-tune training including adjusting, based on the labels, one or more weights of the projection module while maintaining constant weights of the encoder module, wherein the loss function is based on a logit adjustment loss that is based on logits that are adjusted based on a class distribution of the training data. However, Chen teaches: pre-train the artificial neural network to generate matrix representations that are invariant to a predetermined set of data augmentations applied to a sample; and (Chen, paragraph 0084: “In particular, z = g ( h ) is trained to be invariant to data transformations.” – z is analogous to the matrix representations and data transformations is analogous to data augmentations.) after the pre-training, fine-tune train the artificial neural network using a loss function, the fine-tune training including adjusting, based on the labels, one or more weights of the projection module while maintaining constant weights of the encoder module, (Chen, paragraph 0006: “evaluating a loss function that evaluates a difference between the first projected representation and the second projected representation, modifying one or more values of one or more parameters of one or both of the base encoder neural network and the projection head neural network based at least in part on the loss function… performing fine-tuning of the image classification model based on a set of labeled images..” – modifying the parameters of one or both of the base encoder neural network and the projection head neural network indicates that it is capable of adjusting the weights of the projection module while maintaining constant weights to the encoder module while fine-tuning based on a set of labeled images is fine-tuning the artificial neural network based on the labels.) Chen is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and self-supervised contrastive learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li, which already teaches a neural network with an encoder and a projection module configured to generate matrix representations but does not explicitly teach that the matrix representations are invariant or that the fine-tuning includes adjusting weights of only the projection module, to include the teachings of Chen which does teach that the matrix representations are invariant or that the fine-tuning includes adjusting weights of only the projection module in order to "provide improved visual representations." (Chen, abstract) Li and Chen do not explicitly teach: wherein the loss function is based on a logit adjustment loss that is based on logits that are adjusted based on a class distribution of the training data. (Menon, section 5.1: “To do so, consider the following logit adjustment softmax cross-entropy loss for τ > 0 : PNG media_image1.png 84 912 media_image1.png Greyscale Given a scorer that minimizes the above, we now predict a r g m a x y ∈ L f y ( x ) as usual. Compared to the standard softmax cross-entropy (1), the above applies a label-dependent offset to each logit.” – The label-dependent offset indicates that the logit adjustment loss (equation 10) is based on a class distribution.) Menon is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and learning under class imbalance. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li and Chen, which already teaches pre-training and fine-tuning a neural network but does not explicitly teach that the loss function used is based on a logit adjustment loss based on a class distribution of the training data, to include the teachings of Menon which does teach that the loss function used is based on a logit adjustment loss based on a class distribution of the training data as it leads to superior performance. “However, unlike most previous techniques from these streams, logit adjustment is endowed with a clear statistical grounding: by construction, the optimal solution under such adjustment coincides with the Bayes-optimal solution (7) for the balanced error, i.e., it is Fisher consistent for minimising the balanced error. We shall demonstrate this translates into superior empirical performance (§6).” (Menon, section 3, paragraph 3) “However, §3 indicates that our loss in (10) has a firm statistical grounding: it ensures Fisher consistency for the balanced error.” (Menon, section 5.2, paragraph 3) Regarding claim 19, Li, Chen, and Menon teach the system of claim 18, as cited above. Li further teaches: wherein the samples are image samples, and wherein the data augmentations are image transformations. (Li, paragraph 0028: “The neural network 130 may project the training images 110a-n to a low-dimensional subspace, e.g., by generating a “weak” normalized embedding 135 from a weakly augmented image sample and a “strong” normalized embedding 136 from a strongly augmented image sample from the same training image sample.” – The weakly and strongly augmented image samples are analogous to image transformations, see fig. 2, numbers 241 and 242.) Regarding claim 20, Li teaches: A method for performing a task using an artificial neural network fine-tune trained with training data including data samples and corresponding labels, the method comprising: (Li, paragraph 0003: “For example, a neural network classifier may predict a class of the input information among a predetermined set of classes, e.g., identifying that a photo of a fruit orange belongs to the class “orange.” To achieve this, neural networks learn to make predictions gradually, by a process of trial and error, using a machine learning process. A given neural network model may be trained using a large number of training examples, proceeding iteratively until the neural network model begins to consistently make similar inferences from the training examples that a human may make.” – A neural network classifier is a neural network trained to perform a task.) …the artificial neural network including an encoder module followed by a projection module and configured to generate matrix representations based on input samples; and (Li, paragraph 0033: “The neural network 130 includes at least three components: (1) a deep encoder, e.g., a convolutional neural network (CNN) 245 a-c (or collectively referred to as 245) that encodes an image xi or an augmented sample of image xi to a high-dimensional feature vi; (2) a classifier 255 (e.g., a fully-connected layer followed by softmax) that receives vi as input and outputs class predictions; (3) a linear autoencoder 260 a-c (or collectively referred to as 260) that projects vi into a low-dimensional embedding z i ∈ R d .” – The deep encoder is analogous to the encoder module and the autoencoder is analogous to the projection module, while the augmented sample of image xi to a high-dimensional feature vi and the low-dimensional embedding z i are analogous to the matrix representations based on one of the samples.) Li does not explicitly teach: receiving an image by the artificial neural network configured to perform a task based on received images,… processing the image using the artificial neural network to perform the task, wherein the artificial neural network is pre-trained to generate matrix representations that are invariant to a predetermined set of data augmentations applied to received images, and wherein the artificial neural network is, after the pre-training, fine-tune trained using a loss function, the fine-tune training including adjusting, based on the labels, one or more weights of the projection module while maintaining constant weights of the encoder module, wherein the loss function is based on a logit adjustment loss that is based on logits that are adjusted based on a class distribution of the training data. However, Chen teaches: receiving an image by the artificial neural network configured to perform a task based on received images,… processing the image using the artificial neural network to perform the task, (Chen, paragraph 0131: “In an example, unlabeled distillation input data 1302 is provided to classification model 1304 and student network 1314 for processing. Classification model 1304 may be an image classification model or any other type of classification model.” And paragraph 0132: “Further, classification head 1310 may receive and process one or more representations to generate classification output 1312, such as a classification prediction, detection prediction, recognition prediction, segmentation prediction, and/or other types of predictions and prediction tasks.” – With the model being an image classification model, the input data would be an image, therefore, the input data being provided to a classification model that generates a classification prediction is analogous to receiving an image configured to perform a task based on received images.) wherein the artificial neural network is pre-trained to generate matrix representations that are invariant to a predetermined set of data augmentations applied to received images, and (Chen, paragraph 0084: “In particular, z = g ( h ) is trained to be invariant to data transformations.” – z is analogous to the matrix representations and data transformations is analogous to data augmentations.) wherein the artificial neural network is, after the pre-training, fine-tune trained using a loss function, the fine-tune training including adjusting, based on the labels, one or more weights of the projection module while maintaining constant weights of the encoder module, (Chen, paragraph 0006: “evaluating a loss function that evaluates a difference between the first projected representation and the second projected representation, modifying one or more values of one or more parameters of one or both of the base encoder neural network and the projection head neural network based at least in part on the loss function… performing fine-tuning of the image classification model based on a set of labeled images..” – modifying the parameters of one or both of the base encoder neural network and the projection head neural network indicates that it is capable of adjusting the weights of the projection module while maintaining constant weights to the encoder module while fine-tuning based on a set of labeled images is fine-tuning the artificial neural network based on the labels.) Chen is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and self-supervised contrastive learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li, which already teaches a neural network with an encoder and a projection module configured to generate matrix representations but does not explicitly teach that the matrix representations are invariant or that the fine-tuning includes adjusting weights of only the projection module, to include the teachings of Chen which does teach that the matrix representations are invariant or that the fine-tuning includes adjusting weights of only the projection module in order to "provide improved visual representations." (Chen, abstract) Li and Chen do not explicitly teach: wherein the loss function is based on a logit adjustment loss that is based on logits that are adjusted based on a class distribution of the training data. However, Menon teaches: wherein the loss function is based on a logit adjustment loss that is based on logits that are adjusted based on a class distribution of the training data. (Menon, section 5.1: “To do so, consider the following logit adjustment softmax cross-entropy loss for τ > 0 : PNG media_image1.png 84 912 media_image1.png Greyscale Given a scorer that minimizes the above, we now predict a r g m a x y ∈ L f y ( x ) as usual. Compared to the standard softmax cross-entropy (1), the above applies a label-dependent offset to each logit.” – The label-dependent offset indicates that the logit adjustment loss (equation 10) is based on a class distribution.) Menon is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and learning under class imbalance. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li and Chen, which already teaches pre-training and fine-tuning a neural network but does not explicitly teach that the loss function used is based on a logit adjustment loss based on a class distribution of the training data, to include the teachings of Menon which does teach that the loss function used is based on a logit adjustment loss based on a class distribution of the training data as it leads to superior performance. “However, unlike most previous techniques from these streams, logit adjustment is endowed with a clear statistical grounding: by construction, the optimal solution under such adjustment coincides with the Bayes-optimal solution (7) for the balanced error, i.e., it is Fisher consistent for minimising the balanced error. We shall demonstrate this translates into superior empirical performance (§6).” (Menon, section 3, paragraph 3) “However, §3 indicates that our loss in (10) has a firm statistical grounding: it ensures Fisher consistency for the balanced error.” (Menon, section 5.2, paragraph 3) Claims 2 and 4 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Chen in view of Menon in further view of Castells et al. (SuperLoss: A Generic Loss for Robust Curriculum Learning), hereinafter Castells. Castells was cited in the IDS filed on June 1, 2022. Regarding claim 2, Li, Chen, and Menon teach the method of claim 1, as cited above. Li and Chen do not explicitly teach: further comprising curriculum learning based on a difference between the logit adjustment loss and a separation parameter defining an expected logit adjustment loss, wherein the loss function includes a term including a predetermined per-sample confidence parameter. However, Menon teaches: further comprising curriculum learning based on a difference between the logit adjustment loss and a separation parameter defining an expected logit adjustment loss, (Menon, section 5.1: ““For more insight into the loss, consider the following pairwise margin loss PNG media_image2.png 46 526 media_image2.png Greyscale for label weights ay > 0, and pairwise label margins ∆ y y ' representing the desired gap between scores for y and y'. For T = l, our logit adjusted loss (10) corresponds to (11) with ay = 1 and ∆ y y ' = log ⁡ log ⁡ π y ' π y . This demands a larger margin between rare positive ( 1r y ~ 0) and dominant negative (1r y' ~ 1) labels, so that scores for dominant classes do not overwhelm those for rare ones.” – The pairwise label margins representing a desired gap between scores for y and y’ is analogous to the expected logit adjustment loss and is therefore a separation parameter.) Li, Chen, and Menon do not explicitly teach: wherein the loss function includes a term including a predetermined per-sample confidence parameter. However, Castells teaches: wherein the loss function includes a term including a predetermined per-sample confidence parameter. (Castels, Fig. 1: “Our approach consists in appending our SuperLoss on top of any existing loss, without changing anything else in the training procedure.” And section 2.2, paragraph 2: “In contrast to existing confidence-aware losses, it only takes two inputs, namely, the task loss l f x i ,   y i (simplified as l i and referred to as input loss in the following) and a confidence parameter σ i .” – The confidence aware loss is capable of taking any loss function and appending a confidence parameter to it.) Castells is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and curriculum learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li, Chen, and Menon, which already teaches fine-tuning a neural network using a loss function but does not explicitly teach that the loss function includes a term including a predetermined per-sample confidence parameter, to include the teachings of Castells which does teach that the loss function includes a term including a predetermined per-sample confidence parameter in order to "handle difficult samples without restoring to heuristics such as using robust versions of the loss, by instead modulating the loss amplitude w.r.t. the confidence parameter [38]." (Castells, section 2.1, paragraph 2) Regarding claim 4, Li, Chen, and Menon teach the method of claim 1, as cited above. Li, Chen, and Menon do not explicitly teach: wherein the logit adjustment loss is determined based on a softmax over the logits that are adjusted based on the class distribution. However, Castells teaches: wherein the logit adjustment loss is determined based on a softmax over the logits that are adjusted based on the class distribution. (Castells, section 2.4, Object Detection: “RetinaNet classification loss is a class-balanced focal loss (FL): l F L p , y = - α y 1 - p y γ log ⁡ p y with p the probabilities predicted by the network for each box obtained with a softmax on the logits z = f x ,   γ ≥   0 , a focusing hyper-parameter and α y a class-balancing hyper-parameter.” – The loss being a logit adjustment loss is already taught by Menon in claim 1, the class-balancing hyperparameters indicates that the logits are adjusted based on the class distribution.) Castells is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and curriculum learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li, Chen, and Menon, which already teaches fine-tuning a neural network using a loss function but does not explicitly teach that the logit adjustment loss function is based on a softmax that are adjusted based on the class distribution, to include the teachings of Castells which does teach that the logit adjustment loss function is based on a softmax that are adjusted based on the class distribution in order to estimate "confidence of positive and negative detections on the fly from their individual loss." (Castells, section 2.4, Object Detection) Claims 8 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Chen in view of Menon, in further view of Chen et al. (Big Self-Supervised Models are Strong Semi-Supervised Learners), hereinafter T.Chen, in further view of Wen et al. (Time Series Anomaly Detection Using Convolutional Neural Networks and Transfer Learning), hereinafter Wen. T.Chen was cited in the IDS filed on June 1, 2022. Regarding claim 8, Li, Chen, and Menon teach the method of claim 1, as cited above. Li, Chen, and Menon do not explicitly teach: wherein the projection module includes two or more fully-connected layers, wherein adjusting one or more of the weights of the projection module includes adjusting one or more of the weights of at least one of the two or more fully-connected layers. However, T.Chen teaches: wherein the projection module includes two or more fully-connected layers, (T.Chen, section 3.3: “To study the effects of projection head for fine-tuning, we pretrain ResNet-50 using SimCLRv2 with different numbers of projection head layers (from 2 to 4 fully connected layers), and examine performance when fine-tuning from different layers of the projection head.”) T.Chen is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and image classification. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li, Chen, and Menon, which already teaches pre-training and fine-tuning a neural network that has an encoder and projection module but does not explicitly teach that the projection module has two or more fully-connected layers, to include the teachings of T.Chen which does teach that the projection module has two or more fully-connected layers since a deeper projection head improves the representation quality (T.Chen, section 1, page 2, bullet 3) Li, Chen, Menon, and T.Chen do not explicitly teach: wherein adjusting one or more of the weights of the projection module includes adjusting one or more of the weights of at least one of the two or more fully-connected layers. However, Wen teaches: wherein adjusting one or more of the weights of the projection module includes adjusting one or more of the weights of at least one of the two or more fully-connected layers. (Wen, section 4.1: We found two fine-tuning strategies with good performance in our tests. The first one is to set up different learning rate multipliers in 10 sections (5 encoding sections, 4 decoding sections, and the output section) as 0.01, 0.04, 0.09, …, 0.81, 1.0. The other one is to freeze the weights in the first two sections and only fine-tune the subsequent sections, and then to unfreeze the first two sections and fine-tune all weights.” – Freezing the weights in sections and only fine-tuning the others, which would include on or more layers, is analogous to adjusting the weights in one or more layers.) Wen is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li, Chen, Menon, and T.Chen, which already teaches a neural network that is pre-trained and fine-tuned and includes a projection module with two or more fully-connected layers but does not explicitly teach adjusting one or more weights of at least one of the fully- connected layers, to include the teachings of Wen which does teach adjusting one or more weights of at least one of the fully connected layers as it has “robustness and ability to handle complex signals.” (Wen, section 6, paragraph 4) Regarding claim 10, Li, Chen, Menon, T.Chen, and Wen teach the method of claim 8, as cited above. Li further teaches: wherein, when a noise level of the labels is greater than a predetermined value, (Li, paragraph 0079: “The subject technology can utilize the curriculum learned by the label correction process of the subject technology for training on datasets with noise above a certain threshold number of images.”) Li, Chen, and Menon do not explicitly teach: wherein the projection module includes three fully-connected layers, and adjusting one or more of the weights includes adjusting one or more of the weights of only the last one of the two fully-connected layers and maintaining constant weights of the first one of the two fully-connected layers. However, T.Chen teaches: wherein the projection module includes three fully-connected layers, and (T.Chen, section 3.3: “To study the effects of projection head for fine-tuning, we pretrain ResNet-50 using SimCLRv2 with different numbers of projection head layers (from 2 to 4 fully connected layers), and examine performance when fine-tuning from different layers of the projection head.”) Li, Chen, Menon, and T.Chen do not explicitly teach: adjusting one or more of the weights includes adjusting one or more of the weights of only the last one of the two fully-connected layers and maintaining constant weights of the first one of the two fully-connected layers. (Wen, section 4.1: We found two fine-tuning strategies with good performance in our tests. The first one is to set up different learning rate multipliers in 10 sections (5 encoding sections, 4 decoding sections, and the output section) as 0.01, 0.04, 0.09, …, 0.81, 1.0. The other one is to freeze the weights in the first two sections and only fine-tune the subsequent sections, and then to unfreeze the first two sections and fine-tune all weights.” – Freezing the weights in sections and only fine-tuning the others, which would include on or more layers, is analogous to adjusting the weights the last one of the fully-connected layers and maintaining the weights constant in the other.) Claims 14, 17, and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Chen in view of Menon in further view of Zhang et al., (Image-on-Scalar Regression via Deep Neural Networks), hereinafter Zhang. Regarding claim 14, Li, Chen, and Menon teach the method of claim 1, as cited above. Li, Chen, and Menon do not explicitly teach: further comprising, by the artificial neural network, performing image regression after the fine-tune training. However, Zhang teaches: further comprising, by the artificial neural network, performing image regression after the fine-tune training. (Zhang, section 4: “In this work, we have presented the novel method NN-ISR that utilizes neural networks to perform estimation and selection for the spatially varying coefficient function of the main effects in image-on-scalar regression.” – The image-on-scalar regression is analogous to performing image regression.) Zhang is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and image analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li, Chen, and Menon, which already teaches pre-training and fine-tuning a neural network but does not explicitly teach that the trained neural network is used to perform image regression, to include the teachings of Zhang which does teach that the trained neural network is used to perform image regression since it “is flexible enough to approximate a wide range of sparse and piecewise-continuous spatially varying functions.” (Zhang, section 4) Regarding claim 17, Li, Chen, and Menon teach the method of claim 1, as cited above. Li further teaches: wherein the artificial neural network is trained to perform … an image classification task … (Li, paragraph 0003: “For example, a neural network classifier may predict a class of the input information among a predetermined set of classes, e.g., identifying that a photo of a fruit orange belongs to the class “orange.” To achieve this, neural networks learn to make predictions gradually, by a process of trial and error, using a machine learning process. A given neural network model may be trained using a large number of training examples, proceeding iteratively until the neural network model begins to consistently make similar inferences from the training examples that a human may make.” – A neural network classifier that can determine classes of photos is a neural network trained to perform the task of image classification.) Li, Chen, and Menon do not explicitly teach: wherein the artificial neural network is trained to perform …an image regression task. However, Zhang teaches: wherein the artificial neural network is trained to perform …an image regression task. (Zhang, section 4: “In this work, we have presented the novel method NN-ISR that utilizes neural networks to perform estimation and selection for the spatially varying coefficient function of the main effects in image-on-scalar regression.” – The image-on-scalar regression is analogous to performing image regression.) Zhang is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and image analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li, Chen, and Menon, which already teaches pre-training and fine-tuning a neural network but does not explicitly teach that the trained neural network is used to perform image regression, to include the teachings of Zhang which does teach that the trained neural network is used to perform image regression since it “is flexible enough to approximate a wide range of sparse and piecewise-continuous spatially varying functions.” (Zhang, section 4) Regarding claim 21, Li, Chen, and Menon teach the method of claim 20, as cited above. Li further teaches: wherein the task is …image classification … (Li, paragraph 0003: “For example, a neural network classifier may predict a class of the input information among a predetermined set of classes, e.g., identifying that a photo of a fruit orange belongs to the class “orange.” To achieve this, neural networks learn to make predictions gradually, by a process of trial and error, using a machine learning process. A given neural network model may be trained using a large number of training examples, proceeding iteratively until the neural network model begins to consistently make similar inferences from the training examples that a human may make.” – A neural network classifier that can determine classes of photos is a neural network trained to perform the task of image classification.) Li, Chen, and Menon do not explicitly teach: wherein the task is …image regression. However, Zhang teaches: wherein the task is …image regression. (Zhang, section 4: “In this work, we have presented the novel method NN-ISR that utilizes neural networks to perform estimation and selection for the spatially varying coefficient function of the main effects in image-on-scalar regression.” – The image-on-scalar regression is analogous to performing image regression.) Zhang is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and image analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li, Chen, and Menon, which already teaches pre-training and fine-tuning a neural network but does not explicitly teach that the trained neural network is used to perform image regression, to include the teachings of Zhang which does teach that the trained neural network is used to perform image regression since it “is flexible enough to approximate a wide range of sparse and piecewise-continuous spatially varying functions.” (Zhang, section 4) Claims 15 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Chen in view of Menon in further view of Zbontar et al. (Barlow Twins: Self-Supervised Learning via Redundancy Reduction), hereinafter Zbontar. Zbontar was cited in the IDS filed on June 1, 2022. Regarding claim 15, Li, Chen, and Menon teach the method of claim 1, as cited above. Li further teaches: wherein the pre-training includes self- supervised learning based on a contrastive loss for negative and positive pairs of samples constructed from the training data, … (Li, paragraph 0036: “The consistency contrastive loss maximizes the inner product between the pair of positive embeddings z, and z1c,) corresponding to the same source image, while minimizing the inner product between 2(b-I) pairs of negative embeddings corresponding to different images. By mapping different views (augmentations) of the same image to neighboring embeddings, the consistency contrastive loss encourages the neural network 130 to learn discriminative representation that is robust to low-level image corruption.” – The contrastive loss maximizing between the positive embeddings and minimizing between the negative embeddings is analogous to the supervised learning being based on a contrastive loss for negative and positive pairs of samples.) Li, Chen, and Menon do not explicitly teach: wherein the pre-training includes self- supervised learning based on … a self-supervised learning method employing a redundancy reduction loss. However, Zbontar teaches: wherein the pre-training includes self- supervised learning based on … a self-supervised learning method employing a redundancy reduction loss. (Zbontar, section 5: “BARLOW TWINS learns self-supervised representations through a joint embedding of distorted images, with an objective function that maximizes similarity between the embedding vectors while reducing redundancy between their components.” – The objective function maximizing similarity while reducing redundancy is analogous to a redundancy reduction loss.) Zbontar is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and image classification. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li, Chen, and Menon, which already teaches pre-training a neural network but does not explicitly teach that the pre-training includes self-supervised learning employing a redundancy loss, to include the teachings of Zbontar which does teach that the pre-training includes self-supervised learning employing a redundancy loss as it “does not require large batches of samples, nor does it require any particular asymmetry in the twin network structure.” (Zbontar, section 5) Regarding claim 16, Li, Chen, and Menon teach the method of claim 1, as cited above. Li, Chen, and Menon do not explicitly teach: by a prediction module, during the pre-training, generating second matrix representations based on the samples, respectively, wherein the pre-training includes pre-training the artificial neural network and the prediction module based on minimizing a similarity loss determined based on the matrix representations and the second matrix representations. However, Zbontar teaches: by a prediction module, during the pre-training, generating second matrix representations based on the samples, respectively, (Zbontar, section 2.1: “The two batches of distorted views Y A and Y B are then fed into a function f θ , typically a deep network with trainable parameters θ, production batches of embeddings Z A and Z B respectively.” – Two embeddings is analogous to having a first and second matrix representation based on the samples.) wherein the pre-training includes pre-training the artificial neural network and the prediction module based on minimizing a similarity loss determined based on the matrix representations and the second matrix representations. (Zbontar, section 1, paragraph 3: “Based on this principle, we propose an objective function which tries to make the cross-correlation matrix computed from twin embeddings as close to the identity matrix as possible.” – The cross-correlation matrix is used in the loss function (equation 1) and therefore is analogous to minimizing a similarity loss based on the matrix representations and second matrix representations.) Zbontar is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning and image classification. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Li, Chen, and Menon, which already teaches pre-training the neural network but does not explicitly teach that pretraining the neural network and prediction module are based on minimizing a similarity loss, to include the teachings of Zbontar which does teach that pretraining the neural network and prediction module are based on minimizing a similarity loss since it is “conceptually simple, easy to implement and learns useful representations as opposed to trivial solutions.” (Zbontar, section 1, paragraph 3) Allowable Subject Matter Claims 3 and 9 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: Regarding claim 3, the separation parameter as a running average of the logit adjustment loss is not taught or suggested by the prior art of record. The closest prior art is Arik et al. (US20210089870), hereinafter Arik. Arik teaches using a moving average in the loss function (paragraph 0030). While it does discuss a cross-entropy loss (paragraph 0028) it does not teach that the loss function is a logit loss and the architecture of Arik is a reinforcement learning model rather than an encoder module and projection head module. Regarding claim 9, adjusting weights of the middle layer of the three fully-connected layers while maintaining constant weights of the other ones of the three fully-connected layers is not taught or suggested by the prior art of record. The closest prior art is T.Chen and Wen, as cited in the previous office action. While T.Chen teaches that the projection module includes three fully-connected layers, it does not discuss freezing layers. Likewise, while Wen teaches freezing sections of the network, it does not specifically teach freezing the front and back while adjusting weights in the middle layer. Response to Arguments Applicant’s arguments, see page, filed June 12, 206, with respect to claims 3 and 9 have been fully considered and are persuasive. The rejection of November 11, 2025 of claims 3 and 9 has been withdrawn. Applicant's arguments filed June 12, 2026 have been fully considered but they are not persuasive. In response to applicant's argument on page 9 that the examiner's conclusion of obviousness is based upon improper hindsight reasoning, it must be recognized that any judgment on obviousness is in a sense necessarily a reconstruction based upon hindsight reasoning. But so long as it takes into account only knowledge which was within the level of ordinary skill at the time the claimed invention was made, and does not include knowledge gleaned only from the applicant's disclosure, such a reconstruction is proper. See In re McLaughlin, 443 F.2d 1392, 170 USPQ 209 (CCPA 1971). Regard applicant’s argument on page 11 that Chen’s “one or both” statement does not teach or suggest two conditions happening at the same time, examiner disagrees. In order to adjust the weights of only one and not the other would imply adjusting one while holding the other frozen, otherwise just one would not be adjusted alone. Chen’s one or both modification would result in either the encoder being updated while the projection head is frozen, the encoder being frozen while the projection head is modified, or both the encoder and projection head being modified. It would not be unreasonable that a POSITA would experiment and chose the method of freezing the encoder while updating the projection head after finding that this method would further speed up training. Regarding applicant’s argument on page 13 regarding Chen’s loss function, paragraph 0006 discusses fine-tuning based on a set of labeled images (see rejection to claim 1 above). Additionally, Chen contemplates using different loss functions within its framework in paragraph 0110: “For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function). Various loss functions can be used such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and/or various other loss functions such as those contained in FIG. 9.” Likewise, applicant’s argument on page 14 and 15 that the teachings of Li and Chen do not combine with Menon, examiner disagrees. While Li does not mention logits, Chen does and it is tied to the projection head module (Chen, paragraph 0124: “In an example, a three-layer projection head neural network, g(h,)=WC3l(a(WC2la(WC1lh,)) may be used where a is a ReLU non-linearity (bias not shown), for example, instead of using flask (x,)=Wtaskf (x,) to compute the logits of pre-defined classes where W'ask is the weight for an added task-specific linear layer (bias also not shown).”) Chen additionally contemplates different loss functions, including a cross-entropy loss, which a logit loss is. (Chen, paragraph 0110: “For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function). Various loss functions can be used such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and/or various other loss functions such as those contained in FIG. 9.”) Likewise, Li contemplates different loss functions when it discusses the combined loss. (Li, paragraph 0044: “In some embodiment, the prototypical contrastive loss, the consistency contrastive loss, the cross-entropy loss and the reconstruction loss may be combined to jointly train the neural network. The combined loss module 280 may then compute an overall training objective, as to minimize a weighted sum of all losses:”) Thus, the combination of Li, Chen, and Menon teaches using the logit loss on the projection head. Further, examiner disagrees with applicant’s argument that Menon’s loss function is specific to the architecture and training of Menon as it states on page 6, section 5: “We now show how to directly bake logit adjustment into the softmax cross-entropy. We show that this approach has an intuitive relation to existing loss modification techniques.” And page 8, paragraph 1: “More broadly, however, there is value in combining logit adjustment with other techniques.” Thus, Menon itself discusses using the logit loss function with different architectures and training. In response to applicant’s argument that there is no teaching, suggestion, or motivation to combine the references on page 18, the examiner recognizes that obviousness may be established by combining or modifying the teachings of the prior art to produce the claimed invention where there is some teaching, suggestion, or motivation to do so found either in the references themselves or in the knowledge generally available to one of ordinary skill in the art. See In re Fine, 837 F.2d 1071, 5 USPQ2d 1596 (Fed. Cir. 1988), In re Jones, 958 F.2d 347, 21 USPQ2d 1941 (Fed. Cir. 1992), and KSR International Co. v. Teleflex, Inc., 550 U.S. 398, 82 USPQ2d 1385 (2007). In this case, Chen discusses that their methods would improve visual representations while Menon discusses that their logit loss leads to improved performance, both of which have been cited in the claim rejections above. Regarding applicant’s arguments on page 22 that Li, Chen, and Menon do not teach estimating the class distribution, examiner disagrees. Claim 5 specifies that the estimating is before the fine-tune training but does not specify when it is performed. Li, paragraph 0086 above does show that there is a known distribution of the class, in order to be able to separate the OOD samples from in-distribution samples. Li, paragraph 0019 further supports this: “As used herein, the term "out-of-distribution" (OOD) input is used to refer to data samples that do not belong to any known classes.” Regarding applicant’s arguments on page 23 that Menon does not teach the difference between the logit adjustment loss and a separation parameter, examiner disagrees. Equation 11 corresponds to equation 10 which is cited as being the logit adjusted loss. Where ∆ y y ' being the pairwise label margins representing desired gaps between scores for y and y’ are the separation parameter. Applicant further argues against the combination of Menon and Castells, however, examiner notes that the combination of Li, Chen, Menon, and Castells teaches the limitations of claim 2. Li teaches that the method is curriculum learning (Li, paragraph 0079: “The subject technology can utilize the curriculum learned by the label correction process of the subject technology for training on datasets with noise above a certain threshold number of images.”) but not that it’s based on a difference between the logit adjustment loss and a separation parameter defining an expected logit adjustment loss. However, this is taught by Menon, as mentioned above. Castells teaches that their per-sample confidence parameter can be combined with any loss function. Castells, Fig. 1: “Our approach consists in appending our SuperLoss on top of any existing loss, without changing anything else in the training procedure.” Page 2, paragraph 2: “As shown in Figure 1, the SuperLoss is simply plugged on top of the original task loss during training, hence its name. Its role is to monitor the loss of each sample during training and to determine the sample contribution dynamically by applying the core principle of curriculum learning.” And page 9, paragraph 1: “Our approach can be used on top of any loss, and thus applied to various tasks: it basically applies the principles of automatic curriculum learning to any learning problem. The main benefit is that it allows to train models that will perform better, especially in the case where training data are corrupted by noise.” Since Castells states that it can be applied on top of any original task loss during training and that it allows to train models that will perform better, a person would seek to add Castells to their models in order to have better performing models. Regarding applicant’s arguments on page 26 that the loss of Menon is unrelated to the loss of Castells, the examiner disagrees. Section 2.4 of Castells discusses different applications for applying the SuperLoss onto different loss functions. Thus, section 2.4, object detection states: “RetinaNet classification loss is a class-balanced focal loss (FL): f,FL(p, y) = -ay(l - Py)'Ylog(py) with p the probabilities predicted by the network for each box obtained with a softmax on the logits z = f(x), 'Y ~ 0 a focusing hyper-parameter and ay a class-balancing hyper-parameter.” Which is discussing a softmax loss function on the logits which is a cross-entropy loss which is taught by Li and Chen and further taught by Menon in the logit loss. Thus, the softmax on the logits would be in Menon’s adjusted-logit framework rather than the SuperLoss portion described by Castells. Additionally, the rational to combine the teaching of Castells to that of Li, Chen, and Menon has been provided as estimating “the confidence of positive and negative detections on the fly from their individual loss.” (Castells, section 2.4, Object Detection) Applicant argues on page 29 that Li, Chen, Menon, T.Chen, and Wen do not teach or suggest the features of claim 10. However, claim is dependent on claim 8, the two or more fully-connected layers is taught by T.Chen in claim 8 and Applicant has not presented an argument showing that T.Chen does not teach that limitation. Likewise, adjusting one or more of the weights of only the last one of the two fully connected layers and maintaining constant weights of the first one of the two fully-connected layers is taught by Wen in claim 8 and Applicant has not presented arguments that Wen does not teach that limitation. Therefore, this is taught by T.Chen and Wen. Further, Li teaches that the adjusting of weights is based on a noise level of the labels being greater than a predetermined value in the rejection of claim 10, specifically Li, paragraph 0079 states: “The subject technology can utilize the curriculum learned by the label correction process of the subject technology for training on datasets with noise above a certain threshold number of images.” Which is further supported as applying to adjusting the weights further in paragraph 0079: “As the pseudo-labels become more accurate, more samples are used to compute the supervised losses.” Which shows that the label correction process is being carried out on the samples and the losses being computed are being used to train the model. T.Chen teaches that the layers are fully connected (section 3.3: “… different numbers of projection head layers (from 2 to 4 fully connected layers), and examine performance when fine-tuning from different layers of the projection head.” Thus, T.Chen teaches the fully connected layers and also contemplates fine-tuning different combinations of the layers. Wen’s sections are analogous to the layers (Wen, section 3.1: “For a time series with length 1024 and C channels, it is encoded by five sections of convolution layers.”) thus, Wen’s freezing the weights in the first two sections and only fine-tune the subsequent sections is analogous to adjusting the weights of only the last one of the two fully-connected layers and maintaining constant weights of the first one of the two fully-connected layers as T.Chen already teaches the two fully-connected layers and Li teaches that adjusting the weights is based on the noise level threshold. Therefore, claim rejections under 35 U.S.C. §103 of claims 1-2, 4-8, and 9-21 are maintained. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Chen et al. (Multi-Level Contrastive Learning for Few-Shot Problems) Krishnan et al. (US20210326660) Menon et al. (US20230017505) Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACQUELINE MEYER whose telephone number is (703)756-5676. The examiner can normally be reached M-F 8:00 am - 4:30 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at 571-272-4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.C.M./Examiner, Art Unit 2144 /TAMARA T KYLE/Supervisory Patent Examiner, Art Unit 2144
Read full office action

Prosecution Timeline

Show 5 earlier events
Nov 26, 2025
Final Rejection mailed — §103
Feb 25, 2026
Notice of Allowance
Feb 25, 2026
Response after Non-Final Action
Apr 25, 2026
Response after Non-Final Action
May 07, 2026
Response after Non-Final Action
Jun 12, 2026
Response after Non-Final Action
Jul 03, 2026
Response after Non-Final Action
Sep 04, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743602
Automated Variational Inference using Stochastic Models with Irregular Beliefs
3y 8m to grant Granted Sep 22, 2026
Patent 12737637
MACHINE LEARNING MODEL TRAINING WITH ADVERSARIAL LEARNING AND TRIPLET LOSS REGULARIZATION
3y 6m to grant Granted Sep 15, 2026
Patent 12717933
METHOD AND SYSTEM FOR SECURING NEURAL NETWORK MODELS
4y 2m to grant Granted Aug 25, 2026
Patent 12718094
SYSTEM AND METHOD FOR CONTINUAL REFINABLE NETWORK
3y 8m to grant Granted Aug 25, 2026
Patent 12705498
SYSTEMS AND METHODS FOR FEDERATED VALIDATION OF MODELS
3y 8m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
65%
Grant Probability
99%
With Interview (+61.7%)
3y 11m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 23 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month