Prosecution Insights
Last updated: September 25, 2026
Application No. 18/456,258

UNKNOWN-CLASS (OUT-OF-DISTRIBUTION) DATA DETECTION IN MACHINE LEARNING MODELS

Final Rejection §103
Filed
Aug 25, 2023
Priority
Aug 25, 2022 — provisional 63/400,970
Examiner
JUNG, DONG YOON
Art Unit
2145
Tech Center
2100 — Computer Architecture & Software
Assignee
University of Central Florida Research Foundation Inc.
OA Round
2 (Final)
Grant Probability
Favorable
3-4
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
21 currently pending
Career history
6
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103
DETAILED ACTION The action is in response to the original filing on August 25, 2023 and the Remarks and Amendments filed on July 16, 2026. Claims 1-20 are pending and have been considered below. Claims 1, 10 and 19 are independent claims. Claims 1-7, 9-15, 18-20 are amended. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority The present application has a provisional application No. 63/400,970 filed on August 25, 2022. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-5, 7, 8, 10-14, 16, 17 are rejected under 35 U.S.C. 103 as being unpatentable over Zaeemzadeh et al. (Zaeemzadeh), Non-Patent Literature listed in IDS, “Out-of-Distribution Detection Using Union of 1-Dimensional Subspaces”, published on a June 2021 conference, Pages: 10, in view of Ho et al. (Ho), Non-Patent Literature listed in IDS, “Contrastive Learning with Adversarial Examples”, published on 2020, Pages: 13, in further view of Mohseni et al. (Mohseni), Japanese Patent, JP-7725460-B2, Japanese-English Translated Version. As to independent Claim 1, Zaeemzadeh teaches a method of detecting an unknown class data set within a largescale dataset, the method comprising: loading, into a memory of a computing device, a predetermined largescale dataset, wherein the largescale dataset comprises a plurality of in-distribution (ID) class data ( Zaeemzadeh, Pg9457, Left Column, Section 5 Experiments, Lines1-10, "For the image classification task, we train the WideResNet model on CIFAR-10 and CIFAR-100 [16] datasets, which consist of 50,000 images for training and 10,000 images for testing, with an image size of 32 x 32. The testing set is used as the ID testing samples. Similarly to prior work [20, 21, 24], for the OOD testing samples, we use the following datasets: (i) TinyImagenet: The Tiny ImageNet dataset consists of 10,000 test images of size 36x36 belonging to 200 different classes, which are sampled from the original 1,000 classes of ImageNet", wherein trained with CIFAR-10/100 or the TinyImagenet using the model indicates a computing device is being used for the training and is inherently equivalent to loading the dataset into a memory as this ID/OOD dataset is largescale such that it must be stored in a memory of a computing device (hereinafter, any computing components related to the computing device will be this computing device)); calculating, via the at least one processor of the computing device, at least one singular vector for the plurality of ID class data by forming, for each ID class of the plurality of ID class data, a class-specific feature matrix from feature vectors extracted from the plurality of ID class data ( Zaeemzadeh, Pg9454, Right Column, Paragraph2, Subsection: First singular vector as a robust representative, Lines7-9, "Let X_l denote an M x N matrix containing N M-dimensional feature vectors belonging to class l” Pg9455, Figure1, circled section of “Feature vectors for class l”, PNG media_image1.png 421 985 media_image1.png Greyscale , wherein Zaeemzadeh explicitly teaches aggregating N feature vectors of dimension M belonging to a specific in-distribution class l to construct a class-specific feature matrix X_l) applying a singular value decomposition algorithm to the class-specific feature matrix ( Zaeemzadeh, Pg9454, Right Column, Paragraph2, Subsection: First singular vector as a robust representative, Lines9-12, "Furthermore, consider the autocorrelation matrix of the class l defined as C_l = X_l XT_l . Eigenvectors and eigen-values of C_l are the left singular vectors and the square of singular values of X_l, respectively" Pg9455, Figure1, Arrow below of “Feature vectors for class l”, Pg9456, Right Column, Lines1-5, "v1(l) can be computed using the extracted features from training ID samples of class l”, wherein Zaeemzadeh teaches executing SVD (or eigenvector decomposition on autocorrelation matrix C_l) on the class feature matrix X_l to compute its left singular vectors and singular values) selecting a first singular vector representative of the feature vectors corresponding to the ID class ( Zaeemzadeh, Pg9452, Abstract, Lines11-13, "Second, the first singular vector of ID samples belonging to a 1-dimensional subspace can be used as their robust representative" Pg9454, Right Column, Paragraph3, Lines33-37, "Therefore, if the feature vectors belonging to the same class lie on a 1-dimensional subspace, we can use the first singular vector of X_l as a robust representative of the class subspace in the feature space and to reject outliers", wherein Zaeemzadeh discloses selecting the dominant (first) singular vector v1(l) , mentioned above, which aligns with the primary direction containing the vast majority of sample variance, as the singular representative vector for each ID class), and storing the first singular vector for use in an out-of-distribution (OOD) detection test for the ID class ( Zaeemzadeh, Pg9452, Right Column, Paragraph1, Lines10-20, “First, due to compact representation in the feature space, OOD samples are less likely to occupy the same region as the known classes. In other words, a random vector in a high-dimensional space lies on a specific 1-dimensional line with probability 0. Second, we show that the first singular vector of a 1-dimensional subspace is a robust representative of its samples. We exploit these two desirable features and reject samples as OOD, if they occupy the region corresponding to the training samples with probability 0. This region is identified by the set of the first singular vectors of the training classes” Pg9455, Figure1, After SVD arrow, PNG media_image2.png 424 997 media_image2.png Greyscale wherein Zaeemzadeh teaches storing the generated first singular vectors v1(l) so they can be referenced at test time to compute cosine/angular similarity against test feature vectors x_n for OOD detection), thereby establishing at least one ID class associated with the plurality of ID class data ( wherein combining these four steps above calculates the representative singular vector establishing the ID class because forming the class-specific matrix and applying SVD extracts the principal direction of the class data’s variance, while selecting and storing the first singular vector mathematically establishes that vector as the definitive benchmark (subspace) representing and identifying that ID class, rendering it functionally equivalent to the claimed invention.) estimating, during deployment of a machine learning model, an uncertainty score for a test datapoint by calculating, for each ID class, a cosine similarity between a feature vector representing the test datapoint and the first singular vector stored for use in the OOD detection test for the ID class ( Zaeemzadeh, Pg9456, Section4, Lines13-18, PNG media_image3.png 252 563 media_image3.png Greyscale Pg9455, Figure1, Lines4-5, “At test time, the cosine similarity between the test samples and the first singular vector corresponding to each class is used to distinguish between the ID and OOD samples” wherein Zaeemzadeh discloses performing OOD detection at test time (corresponding to during deployment of a machine learning model) by estimating the probability PNG media_image4.png 36 141 media_image4.png Greyscale as an uncertainty score for a test datapoint i_n. To derive the spectral discrepancy PNG media_image5.png 29 28 media_image5.png Greyscale , Zaeemzadeh explicitly computes PNG media_image6.png 70 79 media_image6.png Greyscale , which calculates the cosine similarity (specifically, absolute cosine similarity) between the test datapoint’s feature vector x_n and the stored first singular vector v1(l) for each ID class l); determining, based on a comparison of the uncertainty score to a threshold, whether the test datapoint matches at least one ID class ( Zaeemzadeh, Pg9456, Section4, Lines19-21, “ PNG media_image7.png 24 32 media_image7.png Greyscale is a critical spectral discrepancy and defines the region belonging to the known classes” Pg9456, Right Column, Lines12-14, “ PNG media_image7.png 24 32 media_image7.png Greyscale is the decision parameter, which can be set to achieve a problem-specific precision and/or recall requirements…” Pg9452, Abstract, Lines16-19, “At the test time, employing sampling techniques used for approximate Bayesian inference in deep learning, input samples are detected as OOD if they occupy the region corresponding to the ID samples with probability 0” wherein Zaeemzadeh teaches comparing the calculated uncertainty score (e.g., spectral discrepancy PNG media_image5.png 29 28 media_image5.png Greyscale or probability PNG media_image4.png 36 141 media_image4.png Greyscale ) against a critical spectral discrepancy threshold PNG media_image7.png 24 32 media_image7.png Greyscale , which serves as a decision parameter defining the region belonging to known classes, to determine whether the test datapoint i_n matches an ID class or is rejected as an OOD sample.); However, Zaeemzadeh does not explicitly teach Pre-training using self-supervised adversarial contrastive learning to generate a plurality of feature vectors for the plurality of ID class data, wherein the plurality of ID class data is augmented with the adversarial perturbations Augmented ID class data As mentioned above, Zaeemzadeh teaches about loading the predetermined largescale dataset comprising a plurality of ID class data, but is silent about the [limitation 1] above. From the same field of endeavor, Ho teaches this ( Ho, Pg1, Abstract, Lines6-8, "This paper addresses the problem, by introducing a new family of adversarial examples for contrastive learning and using these examples to define a new adversarial training algorithm for SSL, denoted as CLAE" Pg2, Subsection 2.1 Contrastive learning, Lines1-6, "Contrastive learning has been widely used in the metric learning literature and, more recently, for self-supervised learning (SSL), where it is used to learn an encoder in the pretext training stage. Under the SSL setting, where no labels are available, CL algorithms aim to learn an invariant representation of each image in the training set. This is implemented by minimizing a contrastive loss evaluated on pairs of feature vectors extracted from data augmentations of the image" Pg4, Last Paragraph, Lines1-3, PNG media_image8.png 80 918 media_image8.png Greyscale Pg6, Section4.1, Line1, “Experiments are performed on CIFAR10 [29], CIFAR100 [29] or tinyImagenet…” Pg4, Section3.2, Lines1-5, “In SSL, the dataset is unlabeled, i.e. U = {x_i}; and each example x is mapped into an example pair… CL (contrastive learning) seeks to learn an invariant representation of image x_i by minimizing the risk defined by the loss” wherein Ho teaches a self-supervised learning (SSL) pretext training framework (CLAE) that pre-trains an encoder on a largescale training dataset U = {x_i} (e.g., CIFAR-10, CIFAR100, or ImageNet, the corresponding ID/known class data) to generate feature vectors. Furthermore, Ho explicitly augments each ID training data instances x_i with an adversarial perturbation PNG media_image9.png 25 23 media_image9.png Greyscale to construct perturbed training pairs PNG media_image10.png 27 140 media_image10.png Greyscale , thereby generating feature vectors from ID class data augmented with adversarial perturbations, which corresponds to the [limitation 2] of the augmented ID class data.) Zaeemzadeh and Ho are analogous to the claimed invention as they are from the same field of endeavor of deep neural network representation learning and OOD detection. Therefore, it would have been obvious to a person of ordinary skill in the art to combine, before the effective filing date, the OOD detection framework of Zaeemzadeh with the adversarial contrastive pre-training method of Ho. The motivation to combine is taught by Ho ( Ho, Pg1, Abstract, Lines8-13, “When compared to standard CL, the use of adversarial examples creates more challenging positive pairs and adversarial training produces harder negative pairs by accounting for all images in a batch during the optimization. CLAE is compatible with many CL methods in the literature. Experiments show that it improves the performance of several existing CL baselines on multiple datasets” Pg4, Last Paragraph, Lines3-4, “The rationale is that the use of these pairs in (5) increases the challenge of unsupervised learning, encouraging the learning algorithm to produce a more invariant representation”) A person of ordinary skill in the art would have been motivated to employ Ho’s pre-training as an upstream feature extractor for Zaeemzadeh to produce highly invariant and robust feature vectors from adversarially augmented ID data, which synergistically optimizes the intra-class feature alignment required for forming Zaeemzadeh’s 1-Dimensional subspaces and deriving robust class-representative first singular vectors via SVD. However, both Zaeemzadeh and Ho do not teach about (Highlighted parts) Automatically displaying, the uncertainty score for the test datapoint on a display device associated with the computing device; and based on a determination that the test datapoint matches at least one ID class, labeling, recording, or both, in real-time, the test datapoint as belonging to the at least one ID class; or based on a determination that the test datapoint does not match any ID class, labeling, recording, or both, in real-time, the test datapoint as belonging to a new OOD data category. As mentioned above, Zaeemzadeh teaches about estimating an uncertainty score for a test datapoint and determining, based on a comparison of the uncertainty score to a threshold, whether the test datapoint matches at least one ID class or is rejected as on OOD sample, but is silent about automatically displaying the uncertainty score for the test datapoint on a display device associated with the computing device, and labeling, recording, or both, in real-time, the test datapoint as belonging to the at least one ID class or belonging to a new OOD data category based on the determination. However, in the same field of endeavor, Mohseni teaches this ( Mohseni, Pg8, Paragraph1, Lines2-4, “performing another action at block 510 includes, in response to determining that the input data is OOD at block 508, alerting the driver, performing a failover to a redundant backup system, or gradually reducing vehicle speed” Pg17, Paragraph1, Lines1-8, “In at least one embodiment, one or more of the controllers 1236 may receive input (e.g., represented by input data) from the instrument cluster 1232 of the vehicle 1200 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface ("HMI") display 1234, an audible annunciator, a loudspeaker, and/or via other components of the vehicle 1200. In at least one embodiment, the output may include information such as vehicle speed, speed, time, map data (e.g., a high definition map (not shown in FIG. 12A )), location data (e.g., the location of vehicle 1200 on a map, etc.), direction, locations of other vehicles (e.g., an occupancy grid), information about objects and object conditions sensed by controller 1236” Pg94, Paragraph1, Lines 3-6, “transforms data represented as physical quantities, such as electronic quantities, in the computing system's registers and/or memory into other data similarly represented as physical quantities in the computing system's memory, registers, or other such information storage, transmission, or display device” Pg2, Paragraph5, Lines9-13, “neural network 100 is used during inference to identify OOD input data as an unknown object instead of misclassifying it as in some legacy approaches, thereby increasing safety in some systems that may incorporate neural network 100, such as autonomous vehicles, by providing a method in which further action can be taken based, at least in part, on the input being identified as unknown rather than being misclassified” Pg9, Last Paragraph, Lines5-7, “collecting rejected unsafe scenarios/inputs to augment the in-distribution training set. In at least one embodiment, the rejected unsafe scenarios/inputs are collected from vehicles in use on the road” wherein Mohseni teaches automatically presenting output assessment data and detection results on a display device connected to the computing system. Furthermore, wherein Mohseni explicitly teaches evaluating input data during real-time inference (while operating on the road) to identify, categorize, capture, and record/collect rejected OOD inputs as unknown data categories or in-distribution inputs for system operations and dataset augmentation, which directly corresponds to labeling, recording, or both, in real-time, the test datapoint as belonging to an ID class or a new OOD data category based on the determination.) Zaeemzadeh, Ho, and Mohseni are analogous to the claimed invention as they are from the same field of endeavor of out-of-distribution data detection and uncertainty monitoring in machine learning systems. Therefore, it would have been obvious to a person of ordinary skill in the art to combine, before the effective filing date, the 1-dimensional subspace projection and SVD-based spectral discrepancy uncertainty scoring framework of Zaeemzadeh, the self-supervised adversarial contrastive (CLAE) pre-training for robust feature extraction of Ho with the real-time visual display and automated database labeling/recording control structure of Mohseni. The motivation to combine is taught by Mohseni ( Pg2, Background-Art, Lines1-3, “Handling out-of-distribution input data, i.e., input data that a neural network is not trained to classify, can result in increased classification error rates and can use significant memory, time, or computational resources” Pg2, Paragraph5, Lines6-13, “dentifying OOD input data improves the safety of a system…, by providing a method in which further action can be taken based, at least in part, on the input being identified as unknown rather than being misclassified” Pg8, Lines2-4, “performing another action at block 510 includes, in response to determining that the input data is OOD at block 508, alerting the driver, performing a failover to a redundant backup system, or gradually reducing vehicle speed” ) such that incorporating Ho’s adversarial contrastive pre-training into Zaeemzadeh’s upstream feature extractor maximizes class-wise 1-dimensional subspace alignment to extract highly robust representative singular vectors. This enables Zaeemzadeh’s SVD framework to compute precise, low-false-alarm uncertainty scores, which subsequently provides the high-fidelity input necessary for Mohseni’s downstream visual display to automatically present the uncertainty scores to the user and perform real-time automated labeling and recording of the test data points into ID or OOD database categories without false triggers. As to dependent Claim 2, The combination of Zaeemzadeh, Ho, and Mohseni teaches, as mentioned above, all the limitations of Claim 1. It teaches a method of detecting an unknown class dataset within largescale dataset, comprising pre-training an encoder using self-supervised adversarial contrastive learning on augmented ID data, constructing class-specific feature matrices to calculate and store representative first singular vectors via SVD, estimating an uncertainty score via cosine similarity during deployment, and automatically displaying the score while labeling/recording the test datapoint into an ID class or a new OOD category in real-time based on a threshold comparison. Zaeemzadeh further teaches the method of claim 1, wherein applying the singular value decomposition algorithm comprises selecting, as the first singular vector, a singular vector corresponding to a largest singular value ( Zaeemzadeh, Pg9454, Right Column, Paragraph2, Subsection: First singular vector as a robust representative, Lines9-12, "Furthermore, consider the autocorrelation matrix of the class l defined as C_l = X_l XT_l . Eigenvectors and eigen-values of C_l are the left singular vectors and the square of singular values of X_l, respectively" Pg9454, Lines28-36, “Since the singular values represent the amount of energy concentrated along their corresponding singular vector, if almost all of the energy of the data points in each class is concentrated along its corresponding first singular vector, we will have large PNG media_image11.png 22 23 media_image11.png Greyscale and small PNG media_image12.png 21 17 media_image12.png Greyscale , i >= 2 for all the classes. Therefore, if the feature vectors belonging to the same class lie on a 1-dimensional subspace, we can use the first singular vector of X_l as a robust representative…" Pg9456, Right Column, Lines1-5, "v1(l) can be computed using the extracted features from training ID samples of class l”, , wherein Zaeemzadeh explicitly teaches that the eigenvector PNG media_image12.png 21 17 media_image12.png Greyscale of the autocorrelation matrix C_l = X_l XT_l represents the squared singular values of the class feature matrix X_l, where PNG media_image11.png 22 23 media_image11.png Greyscale denotes the maximum eigenvalue corresponding to the largest singular value. By defining the “first singular vector” v1(l) as the eigenvector associated with this dominant eigenvalue PNG media_image11.png 22 23 media_image11.png Greyscale , which captures almost all energy and variance of the class subspace while remaining least sensitive to perturbations, thus disclosing selecting the singular vector associated with the largest singular value. Consequently, the extraction of v1(l) corresponds to select a singular vector corresponding to a largest singular value.) As to dependent Claim 3, The combination of Zaeemzadeh, Ho, and Mohseni teaches, as mentioned above, all the limitations of Claim 2. It teaches that the selection of the first singular vector comprises the largest singular value. Zaeemzadeh further teaches the method of claim 2, wherein only the first singular vector is stored for use in the OOD detection test for the ID class ( Zaeemzadeh, Pg9452, Right Column, Paragraph1, Lines10-20, “First, due to compact representation in the feature space, OOD samples are less likely to occupy the same region as the known classes. In other words, a random vector in a high-dimensional space lies on a specific 1-dimensional line with probability 0. Second, we show that the first singular vector of a 1-dimensional subspace is a robust representative of its samples. We exploit these two desirable features and reject samples as OOD, if they occupy the region corresponding to the training samples with probability 0. This region is identified by the set of the first singular vectors of the training classes” Pg9454, Lines28-37, “Since the singular values represent the amount of energy concentrated along their corresponding singular vector, if almost all of the energy of the data points in each class is concentrated along its corresponding first singular vector, we will have large PNG media_image11.png 22 23 media_image11.png Greyscale and small PNG media_image12.png 21 17 media_image12.png Greyscale , i >= 2 for all the classes. Therefore, if the feature vectors belonging to the same class lie on a 1-dimensional subspace, we can use the first singular vector of X_l as a robust representative “of the class subspace in the feature space and to reject outliers” Zaeemzadeh, Pg9456, Section4, Lines13-18, PNG media_image3.png 252 563 media_image3.png Greyscale wherein Zaeemzadeh explicitly teaches that because the feature vectors of each known class are embedded onto a 1-dimensional subspace, the class subspace is fully captured by a single dominant axis, namely, the first singular vector PNG media_image13.png 36 34 media_image13.png Greyscale . As shown in Equation2, the OOD detection test measures spectral discrepancy PNG media_image5.png 29 28 media_image5.png Greyscale by taking the cosine similarity of a test sample exclusively against the set of first singular vectors, entirely omitting higher-order singular vectors ( PNG media_image13.png 36 34 media_image13.png Greyscale for i>= 2). Because the classification framework relies solely on these unidimensional representatives to establish the rejection boundary, Zaeemzadeh requires retaining only the first singular vector for each ID class for the OOD test.) As to dependent Claim 4, The combination of Zaeemzadeh, Ho, and Mohseni teaches, as mentioned above, all the limitations of Claim 2. It teaches that the selection of the first singular vector comprises the largest singular value. Zaeemzadeh further teaches the method of claim 2, wherein the first singular vector stored for use in the OOD detection test for each ID class represents a direction along which a majority of the feature vectors extracted from the augmented ID class data corresponding to the ID class are distributed ( Zaeemzadeh, Pg9454, Lines28-37, “Since the singular values represent the amount of energy concentrated along their corresponding singular vector, if almost all of the energy of the data points in each class is concentrated along its corresponding first singular vector, we will have large PNG media_image11.png 22 23 media_image11.png Greyscale and small PNG media_image12.png 21 17 media_image12.png Greyscale , i >= 2 for all the classes. Therefore, if the feature vectors belonging to the same class lie on a 1-dimensional subspace, we can use the first singular vector of X_l as a robust representative “of the class subspace in the feature space and to reject outliers” wherein Zaeemzadeh explicitly teaches that singular values quantify the concentration of feature energy along their associated singular vectors. By constraining the feature vectors of each class to lie on a 1-dimensional subspace, Zaeemzadeh ensures that almost all of the data energy and variance are concentrated along the primary axis corresponding to the maximum eigenvalue PNG media_image11.png 22 23 media_image11.png Greyscale . Because concentrating the vast majority of feature energy along the first singular vector PNG media_image13.png 36 34 media_image13.png Greyscale is mathematically equivalent to aligning the dominant distribution of the class data along that specific vector, rendering it equivalent to the claimed invention of the first singular vector represents the direction along which a majority of the feature vectors are distributed.) As to dependent Claim 5, The combination of Zaeemzadeh, Ho, and Mohseni teaches, as mentioned above, all the limitations of Claim 1. It teaches a method of detecting an unknown class dataset within largescale dataset, comprising pre-training an encoder using self-supervised adversarial contrastive learning on augmented ID data, constructing class-specific feature matrices to calculate and store representative first singular vectors via SVD, estimating an uncertainty score via cosine similarity during deployment, and automatically displaying the score while labeling/recording the test datapoint into an ID class or a new OOD category in real-time based on a threshold comparison. Zaeemzadeh further teaches the method of claim 1, wherein estimating the uncertainty score further comprises measuring, for each ID class, an angular similarity between the feature vector representing the test datapoint and the first singular vector stored for use in the OOD detection test for the ID class ( Zaeemzadeh, Pg9456, Section4, Lines13-18, PNG media_image3.png 252 563 media_image3.png Greyscale Pg9456, Right Column, Lines5-13 “To estimate PNG media_image14.png 34 156 media_image14.png Greyscale , we employ Monte Carlo sampling… PNG media_image7.png 24 32 media_image7.png Greyscale is the decision parameter…” wherein Zaeemzadeh explicitly defines the spectral discrepancy PNG media_image5.png 29 28 media_image5.png Greyscale as the minimum angular distance between the test sample’s feature vector x_n and the primary singular vector PNG media_image13.png 36 34 media_image13.png Greyscale across all known classes. This spectral discrepancy directly measures the angular similarity to evaluate the sample’s spatial deviation from each class’s 1-dimensional subspace, which is subsequently converted into an uncertainty probability PNG media_image14.png 34 156 media_image14.png Greyscale via Monte Carlo sampling. Therefore, Zaeemzadeh’s calculation of angular distance/similarity between test feature vectors and class-specific first singular vectors to determine OOD uncertainty is functionally equivalent to the claimed invention.) As to dependent Claim 7, The combination of Zaeemzadeh, Ho and Mohseni teaches, as discussed above, all the limitations of Claim 1. It teaches a method of detecting an unknown class dataset within largescale dataset, comprising pre-training an encoder using self-supervised adversarial contrastive learning on augmented ID data, constructing class-specific feature matrices to calculate and store representative first singular vectors via SVD, estimating an uncertainty score via cosine similarity during deployment, and automatically displaying the score while labeling/recording the test datapoint into an ID class or a new OOD category in real-time based on a threshold comparison. Zaeemzadeh further teaches the method of claim 1, further comprising fine-tuning, via the at least one processor of the computing device, the at least one singular vector using cross-entropy loss ( Zaeemzadeh, Pg9455, Section3.1, Lines3-5, PNG media_image15.png 96 546 media_image15.png Greyscale Pg9455, Subsection 3.1 Enforcing the Structural Constraints, Paragraph1, Lines18-21, "Therefore the final loss function to be minimized is defined as: PNG media_image16.png 63 368 media_image16.png Greyscale where PNG media_image17.png 23 23 media_image17.png Greyscale is angle between the nth feature vector and the weight vector corresponding to its true label" Pg9455, Figure1, PNG media_image18.png 412 983 media_image18.png Greyscale Pg9455, Right Column, Lines9-12, “In other words, the feature extractor, i.e., the deep neural network, is trained such that it can map each input sample in class l onto a predefined 1-dimensional subspace represented by the direction of w_l” Pg9456, Right Column, Lines1-3, “ PNG media_image13.png 36 34 media_image13.png Greyscale can be computed using the extracted features from training ID samples of class l” wherein Zaeemzadeh discloses optimizing and refining feature representations using cross-entropy loss. Furthermore, Zaeemzadeh discloses training the feature extractor via this cross-entropy loss function so that class feature vectors align along 1-dimensional subspaces to derive the representative singular vector. Under the Broadest Reasonable Interpretation (BRI), “fine-tuning… the at least one singular vector using cross-entropy loss” is reasonably interpreted to encompass fine-tuning or optimizing the underlying neural network feature extractor via cross-entropy loss to constrain and refine the resulting singular vectors PNG media_image13.png 36 34 media_image13.png Greyscale representing the class feature space.) As to dependent Claim 8, The combination of Zaeemzadeh, Ho and Mohseni teaches, as discussed above, all the limitations of Claim 7. It teaches about optimizing and refining feature representations using cross-entropy loss. Zaeemzadeh further teaches the method of claim 7, wherein the at least one singular vector is orthogonal to at least one alternative singular vector ( Zaeemzadeh, Pg9455, Subsection 3.1 Enforcing the Structural Constraints, Right Column, Lines2-6, "Minimum interclass cosine similarity can be enforced by ensuring that w_l are orthogonal to each other. We achieve this by simply initializing the weight matrix with orthonormal vectors, as described in [31], and freezing them during the training" Wherein by initializing and freezing the classification layer weights w_l as orthonormal vectors, Zaeemzadeh constrains the 1-dimensional class subspaces to be mutually orthogonal, which results in the derived first singular vectors of different classes being orthogonal to each other.) As to independent Claim 10, it is a system claim that contains similar limitations of Claim 1 and thus rejected under the same rationale. As to dependent Claim 11, it is a system claim that contains similar limitations of Claim 2 and thus rejected under the same rationale. As to dependent Claim 12, it is a system claim that contains similar limitations of Claim 3 and thus rejected under the same rationale. As to dependent Claim 13, it is a system claim that contains similar limitations of Claim 4 and thus rejected under the same rationale. As to dependent Claim 14, it is a system claim that contains similar limitations of Claim 5 and thus rejected under the same rationale. As to dependent Claim 16, it is a system claim that contains similar limitations of Claim 7 and thus rejected under the same rationale. As to dependent Claim 17, it is a system claim that contains similar limitations of Claim 8 and thus rejected under the same rationale. Claims 6, 15 are rejected under 35 U.S.C. 103 as being unpatentable over Zaeemzadeh, Ho and Mohseni, as discussed above in Claims 1 in further view of Thompson et al. (Thompson) Non-Patent Literature, “HOW TRANSFERABLE ARE FEATURES IN CONVOLUTIONAL NEURAL NETWORK ACOUSTIC MODELS ACROSS LANGUAGES”, published in 2019, Pages: 5. As to dependent Claim 6, The combination of Zaeemzadeh, Ho, and Mohseni teaches, as mentioned above, all the limitations of Claim 1. It teaches a method of detecting an unknown class dataset within largescale dataset, comprising pre-training an encoder using self-supervised adversarial contrastive learning on augmented ID data, constructing class-specific feature matrices to calculate and store representative first singular vectors via SVD, estimating an uncertainty score via cosine similarity during deployment, and automatically displaying the score while labeling/recording the test datapoint into an ID class or a new OOD category in real-time based on a threshold comparison. Zaeemzadeh does not teach the method of claim 1, wherein pre-training the predetermined largescale dataset comprises generating the plurality of feature vectors via an encoder of the machine learning model. In the same field of endeavor, Ho teaches this ( Ho, Pg1 Introduction, Lines3-4, “Self-supervised learning (SSL) [25] aims to alleviate this limitation, by leveraging unlabeled data to define surrogate tasks…” Pg2, Subsection 2.1 Contrastive learning, Lines1-3, "Contrastive learning has been widely used in the metric learning literature and, more recently, for self-supervised learning (SSL), where it is used to learn an encoder in the pretext training stage” Pg6, Section4.1, Line1, “Experiments are performed on CIFAR10 [29], CIFAR100 [29] or tinyImagenet…” Pg4, Section3.2, Lines1-5, “In SSL, the dataset is unlabeled, i.e. U = {x_i}; and each example x is mapped into an example pair… CL (contrastive learning) seeks to learn an invariant representation of image x_i by minimizing the risk defined by the loss” wherein Ho explicitly disclose the pre-training using predetermined largescale dataset to generate the feature vectors via an encoder.) However, Zaeemzadeh, Ho and Mohseni do not teach fine-tuning the encoder while freezing at least one weight of a penultimate layer of the encoder From the same field of endeavor, Thompson teaches this limitation ( Thompson, Pg2828, Section3.1, Lines1-4, “The only models that performed considerably worse than the monolingual baseline models were the transfer networks without finetuning whose surgery occurred at one of the fully connected layers (the penultimate two layers)” Pg2829, Section3.4, Lines3-6, “According to this procedure, layers are successively frozen over the course of training, gradually reducing the number of parameters to be updated until, by the end of training, only the last layer is being updated” wherein Thompson explicitly discloses fine-tuning a pre-trained neural network while freezing the parameters of specific layers, explicitly identifying the fully connected layers as the penultimate layers (“the penultimate two layers”) and evaluating model retraining performance under such weight-freezing conditions.) Zaeemzadeh, Ho, Mohseni and Thompson are analogous to the claimed invention as they are from the same field of endeavor of out-of-distribution data detection and feature space representation in deep neural networks. Therefore, it would have been obvious to a person of ordinary skill in the art to combine, before the effective filing date, the mapping feature representations to 1-dimensional orthogonal subspaces of Zaeemzadeh, the self-supervised adversarial contrastive (CLAE) pre-training for robust feature extraction of Ho, the real-time visual display and automated database labeling/recording control structure of Mohseni with freezing at least one weight of a penultimate layer of the encoder of Thompson. The motivation is taught by Thompson ( Thompson, Pg2830, Left Column, Paragraph1, Lines12-13, “Something about freezing all but the last layer(s) facilitates a 2.7 pp improvement over baseline…” Pg2830, Right Column, Paragraph1, Lines12-13 and Lines23-25 “In this way, making fewer parameter updates actually led to significant performance gains”) such that applying Thompson’s technique of freezing weights in the penultimate layer to the encoder during fine-tuning in the combined framework of the three restricts parameter updates in high-level feature representations, thereby stabilizing gradient updates during optimization, preventing overfitting to downstream fine-tuning data, and preserving the pre-aligned 1-dimensional orthogonal subspace geometry to maximize OOD detection performance. As to dependent Claim 15, it is a system claim that contains similar limitations of Claim 6 and thus rejected under the same rationale. Claims 9, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Zaeemzadeh, Mohseni and Ho, as discussed above in Claims 1 in further view of Guo et al. (Guo), Chinese Patent Application No. CN-112734037-A, published in April 30, 2021. As to dependent Claim 9, The combination of Zaeemzadeh, Ho and Mohseni teaches, as discussed above, all the limitations of Claim 7. It teaches about optimizing and refining feature representations using cross-entropy loss. However, the combination does not teach scaling, via the at least one processor of the computing device, the at least one singular vector with at least one sharpening function, thereby increasing a confidence in the at least one singular vector In the same field of endeavor, Guo teaches this limitation ( Guo, Pg7, Paragraph3, Lines1-5, “In order to enhance the confidence level of the prediction result, the invention uses the method for minimizing entropy. Applying sharpening function, Sharpen(p, T), processing the average value of reference label yr and R prediction classification probability, reducing the entropy of label distribution, so as to ensure that the decision boundary does not pass through the high density area of the edge data distribution” Pg11, Claim6, Lines6-8, “processing the average value of the R prediction classification probability and the reference label based on the sharpening function, obtaining the prediction label for the new data” wherein Guo teaches fine-tuning the model parameters and refining feature representations using a cross entropy loss function wherein this fine-tuning step further incorporates processing the prediction probabilities with a temperature-based sharpening function (Sharpen(p,T)) specifically to enhance the confidence level by reducing the entropy of the label distribution. As mentioned in Claim7, under the BRI, fine-tuning the underlying neural network feature representations using cross-entropy loss encompasses the iterative optimization loop where class feature vectors, and their derived singular vectors, are constrained and updated. Because Guo applies the sharpening function directly within this cross-entropy loss-driven fine-tuning pipeline to boost prediction certainty and shape the resulting vector space, renders it functionally equivalent to the claimed invention.) Zaeemzadeh, Ho, Mohseni and Guo are analogous to the claimed invention as they are from the same field of endeavor of deep neural network representation learning, machine learning model optimization, and OOD detection. Therefore, it would have been obvious to a person of ordinary skill in the art to combine, before the effective filing date, the 1-dimensional subspace projection and SVD-based spectral discrepancy uncertainty scoring framework of Zaeemzadeh, the self-supervised adversarial contrastive (CLAE) pre-training of Ho, and the real-time visual display and automated database labeling/recording control structure of Mohseni with scaling feature representations and derived singular vectors using a temperature-scaled sharpening function during cross-entropy fine-tuning of Guo. The motivation is taught by Guo (Guo, Pg7, Paragraph3, Lines1-5, “In order to enhance the confidence level of the prediction result, the invention uses the method for minimizing entropy. Applying sharpening function, Sharpen(p, T), processing the average value of reference label yr and R prediction classification probability, reducing the entropy of label distribution, so as to ensure that the decision boundary does not pass through the high density area of the edge data distribution”) such that incorporating Guo’s temperature-scaled sharpening function into the cross-entropy fine-tuning pipeline of the combined Zaeemzadeh, Ho, and Mohseni system reduces the entropy of the feature label distribution during parameter optimization, thereby scaling and sharpening the representative singular vectors to ensure decision boundaries avoid high-density edge regions, which directly enhances prediction confidence and improves downstream OOD detection accuracy and real-time database labeling reliability. As to dependent Claim 18, it is a system claim that contains similar limitations of Claim 9 and thus rejected under the same rationale. Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Zaeemzadeh et al. (Zaeemzadeh), Non-Patent Literature listed in IDS, “Out-of-Distribution Detection Using Union of 1-Dimensional Subspaces”, published on a June 2021 conference, Pages: 10, in view of Mohseni et al. (Mohseni), Japanese Patent, JP-7725460-B2, Japanese-English Translated Version. As to independent Claim 19, Zaeemzadeh teaches a method of detecting an unknown class data set within a largescale dataset, the method comprising: executing, via a computing device comprising at least one processor, a machine learning model comprising an encoder ( Zaeemzadeh, Pg9457, Section5, Lines1-2, “we train the WideResNet model…” Pg9455, Figure1, “Deep Neural Net, PNG media_image19.png 411 982 media_image19.png Greyscale wherein Zaeemzadeh explicitly teaches running/training the WideResNet/Deep Neural Net model executing a machine learning model comprising an encoder on a computer device (hereinafter the encoder on a computing device will refer to this)); receiving a predetermined largescale dataset comprising a plurality of in-distribution (ID) class data associated with a plurality of ID classes ( Zaeemzadeh, Pg9457, Left Column, Section 5 Experiments, Lines1-10, "For the image classification task, we train the WideResNet model on CIFAR-10 and CIFAR-100 [16] datasets, which consist of 50,000 images for training and 10,000 images for testing, with an image size of 32 x 32. The testing set is used as the ID testing samples. Similarly to prior work [20, 21, 24], for the OOD testing samples, we use the following datasets: (i) TinyImagenet: The Tiny ImageNet dataset consists of 10,000 test images of size 36x36 belonging to 200 different classes, which are sampled from the original 1,000 classes of ImageNet") extracting, via the encoder, a plurality of feature vectors from the plurality of ID class data ( Zaeemzadeh, Pg9454, Right Column, Paragraph2, Subsection: First singular vector as a robust representative, Lines7-9, "Let X_l denote an M x N matrix containing N M-dimensional feature vectors belonging to class l” Pg9455, Figure1, “Deep Neural Net” outputting “Feature vectors for class l”, PNG media_image1.png 421 985 media_image1.png Greyscale Pg9455, Right Column, Lines9-10, “In other words, the feature extractor, i.e., the deep neural network…” wherein Zaeemzadeh explicitly discloses extracting a plurality of feature vectors from the ID class data via the deep neural network feature extractor (the corresponding encoder)); for each ID class of the plurality of ID classes: forming a class-specific feature matrix from feature vectors of the plurality of feature vectors corresponding to the ID class ( Zaeemzadeh, Pg9454, Right Column, Paragraph2, Subsection: First singular vector as a robust representative, Lines7-9, "Let X_l denote an M x N matrix containing N M-dimensional feature vectors belonging to class l” Pg9455, Figure1, circled section of “Feature vectors for class l”, PNG media_image1.png 421 985 media_image1.png Greyscale wherein Zaeemzadeh explicitly teaches aggregating N feature vectors of dimension M belonging to a specific in-distribution class l to construct a class-specific feature matrix X_l) applying a singular value decomposition algorithm to the class-specific feature matrix ( Zaeemzadeh, Pg9454, Right Column, Paragraph2, Subsection: First singular vector as a robust representative, Lines9-12, "Furthermore, consider the autocorrelation matrix of the class l defined as C_l = X_l XT_l . Eigenvectors and eigen-values of C_l are the left singular vectors and the square of singular values of X_l, respectively" Pg9455, Figure1, Arrow below of “Feature vectors for class l”, Pg9456, Right Column, Lines1-5, "v1(l) can be computed using the extracted features from training ID samples of class l”, wherein Zaeemzadeh teaches executing SVD (or eigenvector decomposition on autocorrelation matrix C_l) on the class feature matrix X_l to compute its left singular vectors and singular values) selecting a first singular vector representative of the feature vectors corresponding to the ID class ( Zaeemzadeh, Pg9452, Abstract, Lines11-13, "Second, the first singular vector of ID samples belonging to a 1-dimensional subspace can be used as their robust representative" Pg9454, Right Column, Paragraph3, Lines33-37, "Therefore, if the feature vectors belonging to the same class lie on a 1-dimensional subspace, we can use the first singular vector of X_l as a robust representative of the class subspace in the feature space and to reject outliers", wherein Zaeemzadeh discloses selecting the dominant (first) singular vector v1(l) , mentioned above, which aligns with the primary direction containing the vast majority of sample variance, as the singular representative vector for each ID class), and storing the first singular vector for use in an out-of-distribution (OOD) detection test for the ID class ( Zaeemzadeh, Pg9452, Right Column, Paragraph1, Lines10-20, “First, due to compact representation in the feature space, OOD samples are less likely to occupy the same region as the known classes. In other words, a random vector in a high-dimensional space lies on a specific 1-dimensional line with probability 0. Second, we show that the first singular vector of a 1-dimensional subspace is a robust representative of its samples. We exploit these two desirable features and reject samples as OOD, if they occupy the region corresponding to the training samples with probability 0. This region is identified by the set of the first singular vectors of the training classes” Pg9455, Figure1, After SVD arrow, PNG media_image2.png 424 997 media_image2.png Greyscale wherein Zaeemzadeh teaches storing the generated first singular vectors v1(l) so they can be referenced at test time to compute cosine/angular similarity against test feature vectors x_n for OOD detection); estimating, during deployment of a machine learning model, an uncertainty score for a test datapoint by calculating, for each ID class, a cosine similarity between a feature vector representing the test datapoint and the first singular vector stored for use in the OOD detection test for the ID class ( Zaeemzadeh, Pg9456, Section4, Lines13-18, PNG media_image3.png 252 563 media_image3.png Greyscale Pg9455, Figure1, Lines4-5, “At test time, the cosine similarity between the test samples and the first singular vector corresponding to each class is used to distinguish between the ID and OOD samples” wherein Zaeemzadeh discloses performing OOD detection at test time (corresponding to during deployment of a machine learning model) by estimating the probability PNG media_image4.png 36 141 media_image4.png Greyscale as an uncertainty score for a test datapoint i_n. To derive the spectral discrepancy PNG media_image5.png 29 28 media_image5.png Greyscale , Zaeemzadeh explicitly computes PNG media_image6.png 70 79 media_image6.png Greyscale , which calculates the cosine similarity (specifically, absolute cosine similarity) between the test datapoint’s feature vector x_n and the stored first singular vector v1(l) for each ID class l); determining, based on a comparison of the uncertainty score to a threshold, whether the test datapoint matches at least one ID class ( Zaeemzadeh, Pg9456, Section4, Lines19-21, “ PNG media_image7.png 24 32 media_image7.png Greyscale is a critical spectral discrepancy and defines the region belonging to the known classes” Pg9456, Right Column, Lines12-14, “ PNG media_image7.png 24 32 media_image7.png Greyscale is the decision parameter, which can be set to achieve a problem-specific precision and/or recall requirements…” Pg9452, Abstract, Lines16-19, “At the test time, employing sampling techniques used for approximate Bayesian inference in deep learning, input samples are detected as OOD if they occupy the region corresponding to the ID samples with probability 0” wherein Zaeemzadeh teaches comparing the calculated uncertainty score (e.g., spectral discrepancy PNG media_image5.png 29 28 media_image5.png Greyscale or probability PNG media_image4.png 36 141 media_image4.png Greyscale ) against a critical spectral discrepancy threshold PNG media_image7.png 24 32 media_image7.png Greyscale , which serves as a decision parameter defining the region belonging to known classes, to determine whether the test datapoint i_n matches an ID class or is rejected as an OOD sample.); However, Zaeemzadeh does not explicitly teach Automatically displaying, the uncertainty score for the test datapoint on a displace device associated with the computing device; and based on a determination that the test datapoint matches at least one ID class, labeling, recording, or both, in real-time, the test datapoint as belonging to the at least one ID class; or based on a determination that the test datapoint does not match any ID class, labeling, recording, or both, in real-time, the test datapoint as belonging to a new OOD data category. As mentioned above, Zaeemzadeh teaches about estimating an uncertainty score for a test datapoint and determining, based on a comparison of the uncertainty score to a threshold, whether the test datapoint matches at least one ID class or is rejected as on OOD sample, but is silent about automatically displaying the uncertainty score for the test datapoint on a display device associated with the computing device, and labeling, recording, or both, in real-time, the test datapoint as belonging to the at least one ID class or belonging to a new OOD data category based on the determination. However, in the same field of endeavor, Mohseni teaches this ( Mohseni, Pg8, Paragraph1, Lines2-4, “performing another action at block 510 includes, in response to determining that the input data is OOD at block 508, alerting the driver, performing a failover to a redundant backup system, or gradually reducing vehicle speed” Pg17, Paragraph1, Lines1-8, “In at least one embodiment, one or more of the controllers 1236 may receive input (e.g., represented by input data) from the instrument cluster 1232 of the vehicle 1200 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface ("HMI") display 1234, an audible annunciator, a loudspeaker, and/or via other components of the vehicle 1200. In at least one embodiment, the output may include information such as vehicle speed, speed, time, map data (e.g., a high definition map (not shown in FIG. 12A )), location data (e.g., the location of vehicle 1200 on a map, etc.), direction, locations of other vehicles (e.g., an occupancy grid), information about objects and object conditions sensed by controller 1236” Pg94, Paragraph1, Lines 3-6, “transforms data represented as physical quantities, such as electronic quantities, in the computing system's registers and/or memory into other data similarly represented as physical quantities in the computing system's memory, registers, or other such information storage, transmission, or display device” Pg2, Paragraph5, Lines9-13, “neural network 100 is used during inference to identify OOD input data as an unknown object instead of misclassifying it as in some legacy approaches, thereby increasing safety in some systems that may incorporate neural network 100, such as autonomous vehicles, by providing a method in which further action can be taken based, at least in part, on the input being identified as unknown rather than being misclassified” Pg9, Last Paragraph, Lines5-7, “collecting rejected unsafe scenarios/inputs to augment the in-distribution training set. In at least one embodiment, the rejected unsafe scenarios/inputs are collected from vehicles in use on the road” wherein Mohseni teaches automatically presenting output assessment data and detection results on a display device connected to the computing system. Furthermore, wherein Mohseni explicitly teaches evaluating input data during real-time inference (while operating on the road) to identify, categorize, capture, and record/collect rejected OOD inputs as unknown data categories or in-distribution inputs for system operations and dataset augmentation, which directly corresponds to labeling, recording, or both, in real-time, the test datapoint as belonging to an ID class or a new OOD data category based on the determination.) Zaeemzadeh and Mohseni are analogous to the claimed invention as they are from the same field of endeavor of out-of-distribution data detection and uncertainty monitoring in machine learning systems. Therefore, it would have been obvious to a person of ordinary skill in the art to combine, before the effective filing date, the 1-dimensional subspace projection and SVD-based spectral discrepancy uncertainty scoring framework of Zaeemzadeh with the real-time visual display and automated database labeling/recording control structure of Mohseni. The motivation to combine is taught by Mohseni ( Pg2, Background-Art, Lines1-3, “Handling out-of-distribution input data, i.e., input data that a neural network is not trained to classify, can result in increased classification error rates and can use significant memory, time, or computational resources” Pg2, Paragraph5, Lines6-13, “dentifying OOD input data improves the safety of a system…, by providing a method in which further action can be taken based, at least in part, on the input being identified as unknown rather than being misclassified” Pg8, Lines2-4, “performing another action at block 510 includes, in response to determining that the input data is OOD at block 508, alerting the driver, performing a failover to a redundant backup system, or gradually reducing vehicle speed” ) such that incorporating Mohseni’s real-time visual display and automated database labeling/recording control structure into Zaeemzadeh’s SVD-based spectral discrepancy framework enables the system to automatically present the calculated uncertainty scores to the user and perform real-time automated labeling and recording of test datapoints into ID or new OOD database categories based on the threshold determination, thereby enhancing operational safety and improving dataset management efficiency without false triggers. Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Zaeemzadeh and Mohseni as discussed above in Claim 1 in further view of Ho et al. (Ho), Non-Patent Literature listed in IDS, “Contrastive Learning with Adversarial Examples”, published in 2020, Pages: 13. As to dependent Claim 20, The combination of Zaeemzadeh and Mohseni teaches, as mentioned above, all the limitations of Claim 19. It teaches about extracting feature vectors via an encoder, forming class feature matrices via SVD to store representative first singular vectors, calculating uncertainty scores via cosine similarity during deployment, and displaying/recording detection results in real-time. However, both Zaeemzadeh and Mohseni do not explicitly teach pretraining, via the at least one processor of the computing device, the predetermined largescale dataset using self-supervised adversarial contrastive learning, wherein the plurality of ID class data is augmented with the adversarial perturbations and wherein the plurality of feature vectors used to form the class-specific feature matrices are extracted, via the encoder, from the plurality of ID class data augmented with the adversarial perturbations As mentioned in Claim 19, Zaeemzadeh teaches about the plurality of feature vectors used to form the class-specific feature matrices are extracted, via the encoder ( Zaeemzadeh, Pg9454, Right Column, Paragraph2, Subsection: First singular vector as a robust representative, Lines7-9, "Let X_l denote an M x N matrix containing N M-dimensional feature vectors belonging to class l” Pg9455, Right Column, Lines9-10, “In other words, the feature extractor, i.e., the deep neural network…” Pg9455, Figure1, circled section of “Feature vectors for class l” outputted by an encoder, PNG media_image1.png 421 985 media_image1.png Greyscale ) but is silent about pre-training using self-supervised adversarial contrastive learning to generate an augmented data. From the same field of endeavor, Ho teaches this ( Ho, Pg1, Abstract, Lines6-8, "This paper addresses the problem, by introducing a new family of adversarial examples for contrastive learning and using these examples to define a new adversarial training algorithm for SSL, denoted as CLAE" Pg2, Subsection 2.1 Contrastive learning, Lines1-6, "Contrastive learning has been widely used in the metric learning literature and, more recently, for self-supervised learning (SSL), where it is used to learn an encoder in the pretext training stage. Under the SSL setting, where no labels are available, CL algorithms aim to learn an invariant representation of each image in the training set. This is implemented by minimizing a contrastive loss evaluated on pairs of feature vectors extracted from data augmentations of the image" Pg4, Last Paragraph, Lines1-3, PNG media_image8.png 80 918 media_image8.png Greyscale Pg4, Section3.2, Lines1-5, “In SSL, the dataset is unlabeled, i.e. U = {x_i}; and each example x is mapped into an example pair… CL (contrastive learning) seeks to learn an invariant representation of image x_i by minimizing the risk defined by the loss” wherein Ho explicitly discloses pre-training an encoder using self-supervised adversarial contrastive learning (CLAE) with adversarial perturbations PNG media_image9.png 25 23 media_image9.png Greyscale to generate invariant feature representations, and extracting feature vectors from the adversarially augmented data pairs for downstream feature matrix alignment.) Zaeemzadeh, Mohseni and Ho are analogous to the claimed invention as they are from the same field of endeavor of out-of-distribution data detection and deep neural network representation learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to combine the 1-dimensional subspace projection and SVD-based spectral discrepancy framework of Zaeemzadeh and the real-time display/recording structure of Mohseni with the self-supervised adversarial contrastive pre-training method of Ho. The motivation is taught by Ho (Ho, Pg4, Last Paragraph, Lines3-4, “The rationale is that the use of these pairs in (5) increases the challenge of unsupervised learning, encouraging the learning algorithm to produce a more invariant representation”) such that employing Ho’s adversarial contrastive pre-training to train Zaeemzadeh’s upstream encoder produces highly invariant feature vectors extracted directly from adversarially augmented ID data, which synergistically optimizes the intra-class feature alignment required for forming Zaeemzadeh’s class-specific feature matrices and deriving robust representative first singular vectors via SVD, thereby enhancing downstream OOD detection precision and real-time database labeling accuracy. Response to Arguments 35 USC § 101 Applicant’s arguments and amendments, filed July 16, 2026, regarding the abstract idea rejections from the previous office action made under 35 U.S.C. 101 have been fully considered and they are persuasive. Therefore, the rejections under 35 U.S.C. 101 have been withdrawn. 35 USC § 103 Applicant’s arguments and amendments, filed July 16, 2026, regarding the rejections from the previous office action made under 35 U.S.C. 103 have been fully considered and see below for the details. In the Remarks, Applicant argues in substance that Claims 1-5, 10-14 Independent Claims 1 and 10 have been amended significantly, altering the scope of the claims. In view of newly added limitations, the examiner has overhauled and re-mapped Claims 1-5 and 10-14 under the updated ground of rejection (Zaeemzadeh in view of Ho, and further view of Mohseni). Please refer to the 35 U.S.C. 103 rejection set forth above for the complete analysis addressing the amended limitations. Applicant argues that there is no teaching, suggestion or motivation to combine the references (see 3rd paragraph of page 15 of Remarks). In response to applicant’s argument that there is no teaching, suggestion, or motivation to combine the references, the examiner recognizes that obviousness may be established by combining or modifying the teachings of the prior art to produce the claimed invention where there is some teaching, suggestion, or motivation to do so found either in the references themselves or in the knowledge generally available to one of ordinary skill in the art. See In re Fine, 837 F.2d 1071, 5 USPQ2d 1596 (Fed. Cir. 1988), In re Jones, 958 F.2d 347, 21 USPQ2d 1941 (Fed. Cir. 1992), and KSR International Co. v. Teleflex, Inc., 550 U.S. 398, 82 USPQ2d 1385 (2007). In this case, incorporating Ho’s adversarial contrastive pre-training into Zaeemzadeh’s upstream feature extractor maximizes class-wise 1-dimensional subspace alignment to extract highly robust representative singular vectors. This enables Zaeemzadeh’s SVD framework to compute precise, low-false-alarm uncertainty scores, which subsequently provides the high-fidelity input necessary for Mohseni’s downstream visual display to automatically present the uncertainty scores to the user and perform real-time automated labeling and recording of the test data points into ID or OOD database categories without false triggers. Applicant argues that the cited references do not establish the claimed integrated training and deployment sequence without impermissible hindsight (see 4th paragraph of page 18 of Remarks). In response to applicant's argument that the examiner's conclusion of obviousness is based upon improper hindsight reasoning, it must be recognized that any judgment on obviousness is in a sense necessarily a reconstruction based upon hindsight reasoning. But so long as it takes into account only knowledge which was within the level of ordinary skill at the time the claimed invention was made, and does not include knowledge gleaned only from the applicant's disclosure, such a reconstruction is proper. See In re McLaughlin, 443 F.2d 1392, 170 USPQ 209 (CCPA 1971). As detailed in the rejection above, one of ordinary skill in the art would have been motivated to combine the references based on the explicit teachings of Ho and Mohseni to enhance feature invariance and real-time database management without relying on Applicant’s specification as a roadmap. Claim 6 and 15 Applicant’s arguments with respect to claims 6 and 15 have been considered but are moot in view of the new ground(s) of rejection (Zaeemzadeh, Ho, Mohseni, Thompson), necessitated by Applicant’s substantial amendment (i.e., fine-tuning the encoder while freezing at least one weight of a penultimate layer of the encoder) to the claims which significantly affected the scope thereof. Please see the rejection for claims 6 and 15 above for further details. Claims 7, 8, 16 and 17 Claims 7-8 and 16-17 have minor or no amendments, but under the new ground of rejection established in their parent claims 1, 10 and 7, 16, they are rejected accordingly. Applicant’s arguments regarding cross-entropy training and orthogonal singular vectors have been fully considered. Under the broadest reasonable interpretation (BRI), fine-tuning feature representations via cross-entropy loss encompasses constraining the feature extractor to refine the resulting singular vectors. Further, initializing weights as orthonormal vectors constrains the resulting singular vectors to be orthogonal. Please refer to the updated 35 U.S.C. 103 rejection above for the complete mapping. Claims 9 and 18 Applicant’s arguments with respect to Claims 9 and 18 have been considered but are moot because the new ground of rejection (Zaeemzadeh, Ho, Mohseni, Guo) necessitated by Applicant’s substantial amendment to the parent Claim 1 which significantly affected the scope thereof. Please see the rejection for claims 9 and 18 above for further details. Claims 19 and 20 Independent Claim 19 and dependent Claim 20 have been significantly amended. The examiner has re-mapped Claim 19 under Zaeemzadeh in view Mohseni, and Claim 20 under Zaeemzadeh in view of Mohseni and Ho. Please refer to the detailed 35 U.S.C. 103 rejections set forth above. Applicant argues that the cited references do not establish the claimed integrated training and deployment sequence without impermissible hindsight (see 4th paragraph of page 18 of Remarks). In response to applicant's argument that the examiner's conclusion of obviousness is based upon improper hindsight reasoning, it must be recognized that any judgment on obviousness is in a sense necessarily a reconstruction based upon hindsight reasoning. But so long as it takes into account only knowledge which was within the level of ordinary skill at the time the claimed invention was made, and does not include knowledge gleaned only from the applicant's disclosure, such a reconstruction is proper. See In re McLaughlin, 443 F.2d 1392, 170 USPQ 209 (CCPA 1971). Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DONG YOON JUNG whose telephone number is (571)270-0198. The examiner can normally be reached 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DONG YOON JUNG/Examiner, Art Unit 2145 /CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145
Read full office action

Prosecution Timeline

Aug 25, 2023
Application Filed
May 28, 2026
Non-Final Rejection mailed — §103
Jul 16, 2026
Response Filed
Sep 11, 2026
Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
Grant Probability
Moderate
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month