Prosecution Insights
Last updated: October 02, 2026
Application No. 18/733,518

METHODS AND SYSTEMS FOR TRAINING ARTIFICIAL INTELLIGENCE-BASED MODELS USING LIMITED LABELED DATA

Non-Final OA §101§103
Filed
Jun 04, 2024
Priority
Jun 05, 2023 — IN 202341038556
Examiner
KWON, JUN
Art Unit
Tech Center
Assignee
Mastercard International Incorporated
OA Round
1 (Non-Final)
41%
Grant Probability
Moderate
1-2
OA Rounds
2y 4m
Est. Remaining
88%
With Interview

Examiner Intelligence

Grants 41% of resolved cases
41%
Career Allowance Rate
32 granted / 78 resolved
-19.0% vs TC avg
Strong +47% interview lift
Without
With
+47.2%
Interview Lift
resolved cases with interview
Typical timeline
4y 8m
Avg Prosecution
32 currently pending
Career history
108
Total Applications
across all art units

Statute-Specific Performance

§101
28.0%
-12.0% vs TC avg
§103
48.5%
+8.5% vs TC avg
§102
9.0%
-31.0% vs TC avg
§112
13.7%
-26.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 78 resolved cases

Office Action

§101 §103
Detailed Action Claims 1-20 are presently pending. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1, Step 1: Claim 1 recites a computer-implemented method comprising: transformer models. Therefore, it is directed to the statutory category of processes. 2A Prong 1: generating, categorical features; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating, determining, determining, generating, 2A Prong 2: A computer-implemented method comprising: accessing, by a server system, a tabular dataset from a database associated with the server system, the tabular dataset comprising tabular data related to a plurality of entities, the tabular data comprising labeled data and unlabeled data; (insignificant extra-solution activity MPEP 2106.05(g)(iii) of gathering statistics) generating, by the server system, a set of labeled features (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, by the server system, a set of unlabeled features (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) determining, by the server system via a first transformer model (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) determining, by the server system via a second transformer model (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, by the server system, a set of concatenated embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, by the server system, a third transformer model based, at least in part, on the set of concatenated embeddings. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – generating a generic machine learning model) The additional elements as disclosed above alone or in combination do not integrate the judicial exception into practical application as they are mere insignificant extra solution activity and combination of generic computer functions that are implemented to perform the disclosed abstract idea above. 2B: A computer-implemented method comprising: accessing, by a server system, a tabular dataset from a database associated with the server system, the tabular dataset comprising tabular data related to a plurality of entities, the tabular data comprising labeled data and unlabeled data; (indicated as an insignificant extra-solution activity MPEP 2106.05(g)(iii) in Step 2A Prong 2. Therfore, it is re-evaluated as well understood, routine and conventional activity MPEP 2106.05(d)(II)(iv) of gathering statistics) generating, by the server system, a set of labeled features (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, by the server system, a set of unlabeled features (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) determining, by the server system via a first transformer model (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) determining, by the server system via a second transformer model (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, by the server system, a set of concatenated embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, by the server system, a third transformer model based, at least in part, on the set of concatenated embeddings. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – generating a generic machine learning model) The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are well, understood, routine and conventional activity as disclosed in combination of generic computer functions and usage of elements that are implemented to perform the disclosed abstract idea above. Regarding claim 2, Step 1: Processes, as above. 2A Prong 1: Incorporates the rejection of claim 1. 2A Prong 2: The computer-implemented method as claimed in claim 1, further comprising: generating, by the server system, the first transformer model based, at least in part, on iteratively performing a first set of self-supervised learning operations till a first criterion and a first set of semi-supervised learning operations till a second criterion are met. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – conventional machine learning model training process) 2B: The computer-implemented method as claimed in claim 1, further comprising: generating, by the server system, the first transformer model based, at least in part, on iteratively performing a first set of self-supervised learning operations till a first criterion and a first set of semi-supervised learning operations till a second criterion are met. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – conventional machine learning model training process) Regarding claim 3, Step 1: Processes, as above. 2A Prong 1: generating a set of distorted unlabeled numerical features based, at least in part, on the set of unlabeled numerical features; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating a set of unlabeled numerical embeddings based, at least in part, on the set of unlabeled numerical features; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating a set of distorted unlabeled numerical embeddings based, at least in part, on the set of distorted unlabeled numerical features; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating, generating, computing a self-supervised loss by comparing the first set of contextual numerical embeddings and the second set of contextual numerical embeddings; (mathematical calculation – instant specification [00120]) 2A Prong 2: The computer-implemented method as claimed in claim 2, wherein the first set of self-supervised learning operations comprises: initializing the first transformer model based, at least in part, on a set of first transformer weights; (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – initializing a generic machine learning model) generating, via the first transformer model, a first set of contextual numerical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, via the first transformer model, a second set of contextual numerical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) fine-tuning the first transformer model based, at least in part, on the self-supervised loss, wherein the fine-tuning comprises adjusting the set of first transformer weights. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – conventional machine learning model training process) 2B: The computer-implemented method as claimed in claim 2, wherein the first set of self-supervised learning operations comprises: initializing the first transformer model based, at least in part, on a set of first transformer weights; (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – initializing a generic machine learning model) generating, via the first transformer model, a first set of contextual numerical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, via the first transformer model, a second set of contextual numerical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) fine-tuning the first transformer model based, at least in part, on the self-supervised loss, wherein the fine-tuning comprises adjusting the set of first transformer weights. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – conventional machine learning model training process) Regarding claim 4, Step 1: Processes, as above. 2A Prong 1: generating a set of labeled numerical embeddings based, at least in part, on the set of labeled numerical features; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating a set of distorted labeled numerical embeddings based, at least in part, on the set of labeled numerical embeddings; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating a set of distorted unlabeled numerical features based, at least in part, on the set of unlabeled numerical features; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating a set of distorted unlabeled numerical embeddings based, at least in part, on the set of distorted unlabeled numerical features; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating, generating, computing a semi-supervised loss by comparing the third set of contextual numerical embeddings and the fourth set of contextual numerical embeddings; (mathematical calculation – instant specification [00123]) 2A Prong 2: The computer-implemented method as claimed in claim 2, wherein the first set of semi-supervised learning operations comprises: initializing the first transformer model based, at least in part, on a set of first transformer weights; (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – initialization of a conventional machine learning model) generating, via the first transformer model, a third set of contextual numerical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, via the first transformer model, a fourth set of contextual numerical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) fine-tuning the first transformer model based, at least in part, on the semi-supervised loss, wherein the fine-tuning comprises adjusting the set of first transformer weights. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – conventional machine learning model training process) 2B: The computer-implemented method as claimed in claim 2, wherein the first set of semi-supervised learning operations comprises: initializing the first transformer model based, at least in part, on a set of first transformer weights; (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – initialization of a conventional machine learning model) generating, via the first transformer model, a third set of contextual numerical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, via the first transformer model, a fourth set of contextual numerical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) fine-tuning the first transformer model based, at least in part, on the semi-supervised loss, wherein the fine-tuning comprises adjusting the set of first transformer weights. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – conventional machine learning model training process) Regarding claim 5, Step 1: Processes, as above. 2A Prong 1: Incorporates the rejection of claim 1. 2A Prong 2: The computer-implemented method as claimed in claim 1, further comprising: generating, by the server system, the second transformer model based, at least in part, on iteratively performing a second set of self-supervised learning operations till a third criterion and a second set of semi-supervised learning operations till a fourth criterion. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – conventional machine learning model training process) 2B: The computer-implemented method as claimed in claim 1, further comprising: generating, by the server system, the second transformer model based, at least in part, on iteratively performing a second set of self-supervised learning operations till a third criterion and a second set of semi-supervised learning operations till a fourth criterion. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – conventional machine learning model training process) Regarding claim 6, Step 1: Processes, as above. 2A Prong 1: generating a set of distorted unlabeled categorical features based, at least in part, on the set of unlabeled categorical features; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating a set of unlabeled categorical embeddings based, at least in part, on the set of unlabeled categorical features; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating a set of distorted unlabeled categorical embeddings based, at least in part, on the set of distorted unlabeled categorical features; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating, generating, computing a self-supervised loss by comparing the first set of contextual categorical embeddings and the second set of contextual categorical embeddings; (mathematical calculation – instant specification [00120]) 2A Prong 2: The computer-implemented method as claimed in claim 5, wherein the second set of self-supervised learning operations comprises: initializing the second transformer model based, at least in part, on a set of second transformer weights; (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – initialization of a conventional machine learning model) generating, via the second transformer model, a first set of contextual categorical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, via the second transformer model, a second set of contextual categorical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) fine-tuning the second transformer model based, at least in part, on the self-supervised loss, wherein the fine-tuning comprises adjusting the set of second transformer weights. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – conventional machine learning model training process) 2B: The computer-implemented method as claimed in claim 5, wherein the second set of self-supervised learning operations comprises: initializing the second transformer model based, at least in part, on a set of second transformer weights; (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – initialization of a conventional machine learning model) generating, via the second transformer model, a first set of contextual categorical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, via the second transformer model, a second set of contextual categorical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) fine-tuning the second transformer model based, at least in part, on the self-supervised loss, wherein the fine-tuning comprises adjusting the set of second transformer weights. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – conventional machine learning model training process) Regarding claim 7, Step 1: Processes, as above. 2A Prong 1: generating a set of labeled categorical embeddings based, at least in part, on the set of labeled categorical features; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating a set of distorted labeled categorical embeddings based, at least in part, on the set of labeled categorical embeddings; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating a set of distorted unlabeled categorical features based, at least in part, on the set of unlabeled categorical features; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating a set of distorted unlabeled categorical embeddings based, at least in part, on the set of distorted unlabeled categorical features; (mental process of evaluation – selecting and gathering data, which can be done with the aid of pen and paper) generating, generating, computing a semi-supervised loss by comparing the third set of contextual categorical embeddings and the fourth set of contextual categorical embeddings; (mathematical calculation – instant specification [00123]) 2A Prong 2: The computer-implemented method as claimed in claim 5, wherein the second set of semi-supervised learning operations comprises: initializing the second transformer model based, at least in part, on a set of second transformer weights; (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – initialization of a conventional machine learning model) generating, via the second transformer model, a third set of contextual categorical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, via the second transformer model, a fourth set of contextual categorical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) fine-tuning the second transformer model based, at least in part, on the semi-supervised loss, wherein the fine-tuning comprises adjusting the set of second transformer weights. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – conventional machine learning model training process) 2B: The computer-implemented method as claimed in claim 5, wherein the second set of semi-supervised learning operations comprises: initializing the second transformer model based, at least in part, on a set of second transformer weights; (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – initialization of a conventional machine learning model) generating, via the second transformer model, a third set of contextual categorical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, via the second transformer model, a fourth set of contextual categorical embeddings (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) fine-tuning the second transformer model based, at least in part, on the semi-supervised loss, wherein the fine-tuning comprises adjusting the set of second transformer weights. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – conventional machine learning model training process) Regarding claim 8, Step 1: Processes, as above. 2A Prong 1: generating, generating, 2A Prong 2: The computer-implemented method as claimed in claim 1, further comprising: receiving, by a server system, an input data sample, the input data sample being a data sample for which a prediction has to be performed for a downstream task; (insignificant extra-solution activity MPEP 2106.05(g)(iii) of gathering statistics) generating, by the server system via the third transformer model, an input embedding (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, by the server system via a prediction head, an outcome prediction (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) 2B: The computer-implemented method as claimed in claim 1, further comprising: receiving, by a server system, an input data sample, the input data sample being a data sample for which a prediction has to be performed for a downstream task; (insignificant extra-solution activity MPEP 2106.05(g)(iii) of gathering statistics) generating, by the server system via the third transformer model, an input embedding (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) generating, by the server system via a prediction head, an outcome prediction (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) Regarding claim 9, Step 1: Processes, as above. 2A Prong 1: Incorporates the rejection of claim 8. 2A Prong 2: The computer-implemented method as claimed in claim 8, wherein the prediction head is selected based, at least in part, on the downstream task. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – selecting which machine learning component to generate the output) 2B: The computer-implemented method as claimed in claim 8, wherein the prediction head is selected based, at least in part, on the downstream task. (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f) – selecting which machine learning component to generate the output) Claim 10 is a system claim which recites the same features as method claim 1 and rejected at least for the same reasons. Additional limitations of claim 10 not addressed in claim 1 are addressed below. Step 1: Claim 10 recites a server system comprising: a communication interface; a memory comprising machine-readable instructions; and a processor. Therefore, it is directed to the statutory category of a machine. 2A Prong 1: same as claim 1 2A Prong 2: A server system, comprising: a communication interface; a memory comprising machine-readable instructions; and a processor communicably coupled to the communication interface and the memory, the processor configured to execute the machine-readable instructions to cause the server system at least in part to: (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) 2B: A server system, comprising: a communication interface; a memory comprising machine-readable instructions; and a processor communicably coupled to the communication interface and the memory, the processor configured to execute the machine-readable instructions to cause the server system at least in part to: (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) Claim 11 is a system claim which recites the same features as method claim 2 and rejected at least for the same reasons. Claim 12 is a system claim which recites the same features as method claim 3 and rejected at least for the same reasons. Claim 13 is a system claim which recites the same features as method claim 4 and rejected at least for the same reasons. Claim 14 is a system claim which recites the same features as method claim 5 and rejected at least for the same reasons. Claim 15 is a system claim which recites the same features as method claim 6 and rejected at least for the same reasons. Claim 16 is a system claim which recites the same features as method claim 7 and rejected at least for the same reasons. Claim 17 is a system claim which recites the same features as method claim 8 and rejected at least for the same reasons. Claim 18 is a system claim which recites the same features as method claim 9 and rejected at least for the same reasons. Claim 19 is a non-transitory computer-readable storage medium claim which recites the same features as method claim 1 and rejected at least for the same reasons. Additional limitations of claim 19 not addressed in claim 1 are addressed below. Step 1: Claim 19 recites a non-transitory computer-readable storage medium comprising computer-executable instructions. Therefore, it is directed to the statutory category of a machine. 2A Prong 1: same as claim 1 2A Prong 2: A non-transitory computer-readable storage medium comprising computer-executable instructions that, when executed by at least a processor of a server system, cause the server system to perform a method comprising: (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) 2B: A non-transitory computer-readable storage medium comprising computer-executable instructions that, when executed by at least a processor of a server system, cause the server system to perform a method comprising: (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) Claim 20 is a non-transitory computer-readable storage medium claim which recites the same features as method claim 1 and rejected at least for the same reasons. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 5-6, 8-12, 14-15 and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Luo et al. (“Network On Network for Tabular Data Classification in Real-world Applications”, 2020, hereinafter ‘Luo’) in view of Kong et al. (“A Transformer-Based Contrastive Semi-Supervised Learning Framework for Automatic Modulation Recognition”, Date of Publication 5 April 2023, hereinafter ‘Kong’) in view of Huang et al. (“TabTransformer: Tabular Data Modeling Using Contextual Embeddings”, 2020, hereinafter ‘Huang’) and further in view of Zhu et al. (“XTab: Cross-table Pretraining for Tabular Transformers”, May 10 2023, hereinafter ‘Zhu’). Regarding claim 1, Luo teaches: A computer-implemented method comprising: accessing, by a server system, a tabular dataset from a database associated with the server system, the tabular dataset comprising tabular data related to a plurality of entities, the tabular data comprising labeled data [Luo, page 2320, left col, lines 5-9] and [[page 2320, right col, lines 11-14] show that the method is performed using a conventional computing device with Tensorflow and PyTorch. [page 2319, left col, 2.3 Deep methods, lines 1-17] shows that the input data is tabular data with categorical field features and numerical field features. [3.2 DNN with auxiliary losses, lines 8-9] discloses that the dataset includes label of an instance) determining, by the server system via a first [Luo, page 2319, right col, 3.1 The structure of NON, line 1 – page 2320, left col, line 16] and [page 2320, left col, Figure 3] The field-wise network receives categorical embeddings and numerical embeddings e i to generate the output e i ' = D N N i ( e i ) . Each field-wise neural network are the first neural network and the second neural network. [page 2320, left col, 3.1.2 Across field network, line 1 – right col, line 39] shows that the output generated by the Field-wise network are embeddings as the across field network receives embeddings e j   and features x j ) determining, by the server system via a second [Luo, page 2319, right col, 3.1 The structure of NON, line 1 – page 2320, left col, line 16] and [page 2320, left col, Figure 3] The field-wise network receives categorical embeddings and numerical embeddings e i to generate the output e i ' = D N N i ( e i ) . Each field-wise neural network are the first neural network and the second neural network. [page 2320, left col, 3.1.2 Across field network, line 1 – right col, line 39] shows that the output generated by the Field-wise network are embeddings as the across field network receives embeddings e j   and features x j ) generating, by the server system, a set of concatenated embeddings based, at least in part, on concatenating the set of contextual numerical embeddings and the set of contextual categorical embeddings; and ([page 2320, left col, 3.1.2 Across field network, line 1 – right col, line 39] shows that the output generated by the Field-wise network are embeddings as the across field network receives embeddings e j   and features x j . [page 2320, right col, 3.1.3 Operation fusion network, lines 1-12] generates a final prediction by concatenating the gathered embeddings from the Across field network) However, Luo does not specifically disclose: the tabular data comprising labeled data and unlabeled data; generating, by the server system, a set of labeled features based, at least in part, on the labeled data, the set of labeled features comprising a set of labeled numerical features and a set of labeled categorical features; generating, by the server system, a set of unlabeled features based, at least in part, on the unlabeled data, the set of unlabeled features comprising a set of unlabeled numerical features and a set of unlabeled categorical features; determining, by the server system via a first transformer model, a set of contextual numerical embeddings determining, by the server system via a second transformer model, a set of contextual categorical embeddings generating, by the server system, a third transformer model based, at least in part, on the set of concatenated embeddings. Kong teaches: the [Algorithm 1, lines 1-3] The input training data includes labeled training data set, unlabeled training set, and hyperparameters) generating, by the server system, a set of labeled features based, at least in part, on the labeled data, the set of labeled features comprising a set of labeled [Kong, page 957, Algorithm 1, lines 20-24] The labelled features e i are generated based on labelled dataset D l d . Utilizing a pair of neural networks that processes numerical features and categorical features is taught by the primary reference Luo) generating, by the server system, a set of unlabeled features based, at least in part, on the unlabeled data, the set of unlabeled features comprising a set of unlabeled [Kong, page 957, Algorithm 1, lines 8-15] The unlabeled features e i are generated based on the unlabeled dataset D u d . Utilizing a pair of neural networks that processes numerical features and categorical features is taught by the primary reference Luo) Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Kong and Luo to use the method of utilizing transformer neural networks that receives both labeled data and unlabeled data to of Kong to implement the tabular data machine learning method of the present invention. The suggestion and/or motivation for doing so is to improve the accuracy of the tabular data processing system of Luo by substituting the DNNs to transformers of Kong which can efficiently process large number of datasets and can analyze complex feature interactions, and using loss functions including contrastive loss of Kong can improve the quality of learned representatives significantly [Kong, page 954, left col, 2) Projection Head, lines 1-8]. However, Luo in view of Kong do not specifically disclose: determining, by the server system via a first transformer model, a set of contextual determining, by the server system via a second transformer model, a set of contextual generating, by the server system, a third transformer model based, at least in part, on the set of concatenated embeddings. Huang teaches: determining, by the server system via a first transformer model, a set of contextual [Huang, page 2, right col, 2 The TabTransformer, line 1 – page 3, left col, line 10] The transformer receives parametric embedding and generates contextual embedding. Receiving numerical and categorical embedding is taught by following reference Zhu, and the first and the second neural network is taught by Luo) determining, by the server system via a second transformer model, a set of contextual [Huang, page 2, right col, 2 The TabTransformer, line 1 – page 3, left col, line 10] The transformer receives parametric embedding and generates contextual embedding. Receiving numerical and categorical embedding is taught by following reference Zhu, and the first and the second neural network is taught by Luo) Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Luo, Kong and Huang to use the method of generating contextual embeddings using the transformers of Huang to implement the tabular data machine learning method of the present invention. The suggestion and/or motivation for doing so is to improve the accuracy of the tabular data processing system of Luo [Huang, page 1, Abstract, lines 8-16] and [page 6, right col, last paragraph]. However, Luo in view of Kong and further in view of Huang do not specifically disclose: generating, by the server system, a third transformer model based, at least in part, on the set of concatenated embeddings. Zhu teaches: generating, by the server system, a third transformer model based, at least in part, on the set of concatenated embeddings. ([Zhu, page 3, Figure 1], [page 3, left col, 3.1.1. FEATURIZERS, line 1 – right col, line 25] During the forward pass, the feature embeddings from the input embedding matrices are retrieved using the backbone transformer. The backbone transformer receives all the datasets including the numerical column and the categorical column) Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Luo, Kong, Huang and Zhu to use the method of training a new transformer using both numerical embeddings and categorical embeddings of Zhu to implement the tabular data machine learning method of the instant invention. The suggestion and/or motivation for doing so is to improve the task performance of the tabular data processing system [page 6, left col, 4.3. Comparison with baseline transformers, line 1 – right col, Figure 3]. Regarding claim 2, Luo in view of Kong teaches: The computer-implemented method as claimed in claim 1, generating, by the server system, the first transformer model further comprising: generating, by the server system, the first transformer model based, at least in part, on iteratively performing a first set of self-supervised learning operations till a first criterion and a first set of semi-supervised learning operations till a second criterion are met. ([Kong, Algorithm 1] and [page 957, left col, the first paragraph] TessAMR Update Algorithm has two separate for loops from lines 7 to 17 and from lines 18 to 26. TessAMR Update Algorithm first initializes the parameters of encoder and projection head of a transformer, repeat updating networks until the sampled minibatch from unlabeled dataset is fully used. And then, Kong updates networks using the labeled dataset to further update both f and g. The for loop condition from lines 7 to 17 is the first criterion, and the for loop condition from lines 18 to 26 is the second criterion. [page 955, B. Transformer-Based Encoder] discloses that the f is a transformer) Regarding claim 3, Luo teaches: The computer-implemented method as claimed in claim 2, generating, via the first [Luo, page 2319, right col, 3.1 The structure of NON, line 1 – page 2320, left col, line 16] and [page 2320, left col, Figure 3] The field-wise network receives categorical embeddings and numerical embeddings e i to generate the output e i ' = D N N i ( e i ) . Each field-wise neural network are the first neural network and the second neural network. [page 2320, left col, 3.1.2 Across field network, line 1 – right col, line 39] shows that the output generated by the Field-wise network are embeddings as the across field network receives embeddings e j   and features x j ) However, Luo does not specifically disclose: wherein the first set of self-supervised learning operations comprises: initializing the first transformer model based, at least in part, on a set of first transformer weights; generating a set of distorted unlabeled numerical features based, at least in part, on the set of unlabeled numerical features; generating a set of unlabeled numerical embeddings based, at least in part, on the set of unlabeled numerical features; generating a set of distorted unlabeled numerical embeddings based, at least in part, on the set of distorted unlabeled numerical features; generating, via the first transformer model, a first set of contextual numerical embeddings based, at least in part, on the set of unlabeled numerical embeddings; generating, via the first transformer model, a second set of contextual numerical embeddings based, at least in part, on the set of distorted unlabeled numerical embeddings; computing a self-supervised loss by comparing the first set of contextual numerical embeddings and the second set of contextual numerical embeddings; and fine-tuning the first transformer model based, at least in part, on the self-supervised loss, wherein the fine-tuning comprises adjusting the set of first transformer weights. Kong teaches: wherein the first set of self-supervised learning operations comprises: initializing the first transformer model based, at least in part, on a set of first transformer weights; ([Kong, page 957, Algorithm 1, line 1-6] The parameters of the transformer encoder f and projection head g are initialized) generating a set of distorted unlabeled numerical features based, at least in part, on the set of unlabeled numerical features; ([Kong, page 957, Algorithm 1, lines 8-9] and [page 954, left col, 1) Time Warping, lines 1-3 and lines 27-33] discloses adding noise (Time warping) to obtain the distorted unlabeled dataset x i ^ i = 1 B u ) generating a set of unlabeled numerical embeddings based, at least in part, on the set of unlabeled numerical features; ([Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating embedding e i using the encoder f θ ( x i ) . The Transformer based encoder includes convolutional layers that generates intermediate embeddings when the signal passes through the convolutional layers) generating a set of distorted unlabeled numerical embeddings based, at least in part, on the set of distorted unlabeled numerical features; ([Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating distorted embedding e i ^ using the encoder f θ ( x i ) ^   . The Transformer based encoder includes convolutional layers that generates intermediate embeddings when the signal passes through the convolutional layers) generating, via the first transformer model, a first set of contextual numerical embeddings based, at least in part, on the set of unlabeled numerical embeddings; ([Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating embedding e i using the encoder f θ ( x i ) ) generating, via the first transformer model, a second set of contextual numerical embeddings based, at least in part, on the set of distorted unlabeled numerical embeddings; ([Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating distorted embedding e i ^ using the encoder f θ ( x i ) ^ ) computing a self-supervised loss by comparing the first set of contextual numerical embeddings and the second set of contextual numerical embeddings; and ([Kong, page 957, Algorithm 1, line 13-14] and [page 954, left col, last paragraph, 3) Contrastive Loss, line 1 – right col, line 4] Update networks based on the calculated contrastive loss function. The contrastive loss is calculated using the outputs of the transformer encoder and the projection head, and the based on the similarity between the distorted output v i ^ and normal output v i ^ ) fine-tuning the first transformer model based, at least in part, on the self-supervised loss, wherein the fine-tuning comprises adjusting the set of first transformer weights. ([Kong, page 957, Algorithm 1, line 13-14] and [page 954, left col, last paragraph, 3) Contrastive Loss, line 1 – right col, line 4] Update networks based on the calculated contrastive loss function) Regarding claim 5, Luo in view of Kong teaches: The computer-implemented method as claimed in claim 1, further comprising: generating, by the server system, the second transformer model further comprising: generating, by the server system, the second transformer model based, at least in part, on iteratively performing a second set of self-supervised learning operations till a third criterion and a second set of semi-supervised learning operations till a fourth criterion. ([Kong, Algorithm 1] and [page 957, left col, the first paragraph] TessAMR Update Algorithm has two separate for loops from lines 7 to 17 and from lines 18 to 26. TessAMR Update Algorithm first initializes the parameters of encoder and projection head of a transformer, repeat updating networks until the sampled minibatch from unlabeled dataset is fully used. And then, Kong updates networks using the labeled dataset to further update both f and g. The for loop condition from lines 7 to 17 is the third criterion, and the for loop condition from lines 18 to 26 is the fourth criterion. [page 955, B. Transformer-Based Encoder] discloses that the f is a transformer) Regarding claim 6, Luo teaches: The computer-implemented method as claimed in claim 5, generating, via the second transformer model, a first set of contextual categorical embeddings based, at least in part, on the set of [Luo, page 2319, right col, 3.1 The structure of NON, line 1 – page 2320, left col, line 16] and [page 2320, left col, Figure 3] The field-wise network receives categorical embeddings and numerical embeddings e i to generate the output e i ' = D N N i ( e i ) . Each field-wise neural network are the first neural network and the second neural network. [page 2320, left col, 3.1.2 Across field network, line 1 – right col, line 39] shows that the output generated by the Field-wise network are embeddings as the across field network receives embeddings e j   and features x j ) However, Luo do not specifically disclose: wherein the second set of self-supervised learning operations comprises: initializing the second transformer model based, at least in part, on a set of second transformer weights; generating a set of distorted unlabeled categorical features based, at least in part, on the set of unlabeled categorical features; generating a set of unlabeled categorical embeddings based, at least in part, on the set of unlabeled categorical features; generating a set of distorted unlabeled categorical embeddings based, at least in part, on the set of distorted unlabeled categorical features; generating, via the second transformer model, a first set of contextual categorical embeddings based, at least in part, on the set of unlabeled categorical embeddings; generating, via the second transformer model, a second set of contextual categorical embeddings based, at least in part, on the set of distorted unlabeled categorical embeddings; computing a self-supervised loss by comparing the first set of contextual categorical embeddings and the second set of contextual categorical embeddings; and fine-tuning the second transformer model based, at least in part, on the self-supervised loss, wherein the fine-tuning comprises adjusting the set of second transformer weights. Kong teaches: wherein the second set of self-supervised learning operations comprises: initializing the second transformer model based, at least in part, on a set of second transformer weights; ([Kong, page 957, Algorithm 1, line 1-6] The parameters of the transformer encoder f and projection head g are initialized) generating a set of distorted unlabeled categorical features based, at least in part, on the set of unlabeled categorical features; ([Kong, page 957, Algorithm 1, lines 8-9] and [page 954, left col, 1) Time Warping, lines 1-3 and lines 27-33] discloses adding noise (Time warping) to obtain the distorted unlabeled dataset x i ^ i = 1 B u ) generating a set of unlabeled categorical embeddings based, at least in part, on the set of unlabeled categorical features; ([Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating embedding e i using the encoder f θ ( x i ) . The Transformer based encoder includes convolutional layers that generates intermediate embeddings when the signal passes through the convolutional layers) generating a set of distorted unlabeled categorical embeddings based, at least in part, on the set of distorted unlabeled categorical features; ([Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating distorted embedding e i ^ using the encoder f θ ( x i ) ^   . The Transformer based encoder includes convolutional layers that generates intermediate embeddings when the signal passes through the convolutional layers) generating, via the second transformer model, a first set of contextual categorical embeddings based, at least in part, on the set of unlabeled categorical embeddings; ([Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating embedding e i using the encoder f θ ( x i ) ) generating, via the second transformer model, a second set of contextual categorical embeddings based, at least in part, on the set of distorted unlabeled categorical embeddings; ([Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating distorted embedding e i ^ using the encoder f θ ( x i ) ^ ) computing a self-supervised loss by comparing the first set of contextual categorical embeddings and the second set of contextual categorical embeddings; and ([Kong, page 957, Algorithm 1, line 13-14] and [page 954, left col, last paragraph, 3) Contrastive Loss, line 1 – right col, line 4] Update networks based on the calculated contrastive loss function. The contrastive loss is calculated using the outputs of the transformer encoder and the projection head, and the based on the similarity between the distorted output v i ^ and normal output v i ^ ) fine-tuning the second transformer model based, at least in part, on the self-supervised loss, wherein the fine-tuning comprises adjusting the set of second transformer weights. ([Kong, page 957, Algorithm 1, line 13-14] and [page 954, left col, last paragraph, 3) Contrastive Loss, line 1 – right col, line 4] Update networks based on the calculated contrastive loss function) Regarding claim 8, Luo in view of Kong and further in view of Zhu teaches: The computer-implemented method as claimed in claim 1, further comprising: receiving, by a server system, an input data sample, the input data sample being a data sample for which a prediction has to be performed for a downstream task; ([Zhu, page 3, Figure 1] and [page 3, left col, lines 10-24] The shared backbone transformer is kept for downstream tasks. The input training data is downstream training data and trained until a stopping criterion is met) generating, by the server system via the third transformer model, an input embedding based, at least in part, on the input data sample; and ([Zhu, page 3, Figure 1], [page 3, left col, 3.1.1. FEATURIZERS, line 1 – right col, line 25] During the forward pass, the feature embeddings from the input embedding matrices are retrieved using the backbone transformer. The backbone transformer receives all the datasets including the numerical column and the categorical column) generating, by the server system via a prediction head, an outcome prediction for the downstream task based, at least in part, on the input embedding. ([Zhu, page 3, Figure 1], [page 3, left col, lines 10-24] and [page 3, left col, 3.1.1. FEATURIZERS, line 1 – right col, line 25] The shared backbone transformer is kept for downstream tasks. The input training data is downstream training data and trained until a stopping criterion is met) Regarding claim 9, Luo in view of Kong and further in view of Zhu teaches: The computer-implemented method as claimed in claim 8, wherein the prediction head is selected based, at least in part, on the downstream task. ([Zhu, page 3, Figure 1], [page 3, left col, lines 10-24] and [page 3, left col, 3.1.1. FEATURIZERS, line 1 – right col, line 25] The shared backbone transformer is kept for downstream tasks. The input training data is downstream training data and trained until a stopping criterion is met. [Zhu, page 4, left col, Supervised loss, lines 8-14] The projection heads are data-specific, which means that each projection head process different input data (selected based on the task)) Claim 10 is a system claim which recites the same features as method claim 1 and rejected at least for the same reasons. Additional limitations of claim 10 not addressed in claim 1 are addressed below. Luo teaches: A server system, comprising: a communication interface; a memory comprising machine-readable instructions; and a processor communicably coupled to the communication interface and the memory, the processor configured to execute the machine-readable instructions to cause the server system at least in part to: ([Luo, page 2320, left col, lines 5-9] and [[page 2320, right col, lines 11-14] show that the method is performed using a conventional computing device with Tensorflow and PyTorch) Claim 11 is a system claim which recites the same features as method claim 2 and rejected at least for the same reasons. Claim 12 is a system claim which recites the same features as method claim 3 and rejected at least for the same reasons. Claim 14 is a system claim which recites the same features as method claim 5 and rejected at least for the same reasons. Claim 15 is a system claim which recites the same features as method claim 6 and rejected at least for the same reasons. Claim 17 is a system claim which recites the same features as method claim 8 and rejected at least for the same reasons. Claim 18 is a system claim which recites the same features as method claim 9 and rejected at least for the same reasons. Claim 19 is a non-transitory computer-readable storage medium claim which recites the same features as method claim 1 and rejected at least for the same reasons. Additional limitations of claim 19 not addressed in claim 1 are addressed below. Luo teaches: A non-transitory computer-readable storage medium comprising computer-executable instructions that, when executed by at least a processor of a server system, cause the server system to perform a method comprising: ([Luo, page 2320, left col, lines 5-9] and [[page 2320, right col, lines 11-14] show that the method is performed using a conventional computing device with Tensorflow and PyTorch) Claim 20 is a non-transitory computer-readable storage medium claim which recites the same features as method claim 1 and rejected at least for the same reasons. Claims 4, 7, 13 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Luo in view of Kong in view of Huang in view of Zhu and further in view of Li et al. (US 20210089883 A1, hereinafter ‘Li’). Regarding claim 4, Luo in view of Kong teaches: The computer-implemented method as claimed in claim 2, wherein the first set of semi-supervised learning operations comprises: initializing the first transformer model based, at least in part, on a set of first transformer weights; ([Kong, page 957, Algorithm 1, line 1-6] The parameters of the transformer encoder f and projection head g are initialized. The first and second neural network is taught by Luo) generating a set of labeled numerical embeddings based, at least in part, on the set of labeled numerical features; ([Kong, page 957, Algorithm 1, lines 18-26] Output embedding e i is generated using the encoder transformer f using the labeled dataset D l d . The first neural network processing categorical features and the second neural network processing numerical features is taught by Luo) generating a set of distorted unlabeled numerical features based, at least in part, on the set of unlabeled numerical features; ([Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating distorted feature by applying time warping augmentation and generating an output embedding using the encoder f θ ( x i ) ^   . The Transformer based encoder includes convolutional layers that generates intermediate embeddings when the signal passes through the convolutional layers) generating a set of distorted unlabeled numerical embeddings based, at least in part, on the set of distorted unlabeled numerical features; ([Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating distorted embedding e i ^ using the encoder f θ ( x i ) ^   . The Transformer based encoder includes convolutional layers that generates intermediate embeddings when the signal passes through the convolutional layers) generating, via the first transformer model, a third set of contextual numerical embeddings based, at least in part, on the set of labeled numerical embeddings and the set of unlabeled numerical embeddings; ([Kong, page 957, Algorithm 1, lines 18-26] Output embedding e i is generated using the encoder transformer f using the labeled dataset D l d . [Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating embedding e t using the encoder f θ ( x t )   . The Transformer based encoder includes convolutional layers that generates intermediate embeddings when the signal passes through the convolutional layers) generating, via the first transformer model, a fourth set of contextual numerical embeddings[Kong, page 957, Algorithm 1, lines 18-26] Output embedding e i is generated using the encoder transformer f using the labeled dataset D l d . [Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating distorted embedding e i ^ using the encoder f θ ( x i ) ^   . The Transformer based encoder includes convolutional layers that generates intermediate embeddings when the signal passes through the convolutional layers) computing a semi-supervised loss by comparing the third set of contextual numerical embeddings and the fourth set of contextual numerical embeddings; and ([Kong, page 957, Algorithm 1, lines 18-26] and [page 953, right col, lines 1-10] Calculate the cross entropy loss function, and update networks f and g based on the semi-supervised (total) entropy loss function calculated based on a weighted sum of the unsupervised loss and supervised loss) fine-tuning the first transformer model based, at least in part, on the semi-supervised loss, wherein the fine-tuning comprises adjusting the set of first transformer weights. ([Kong, page 957, Algorithm 1, lines 18-26] and [page 953, right col, lines 1-10] Calculate the cross entropy loss function, and update networks f and g based on the semi-supervised (total) entropy loss function calculated based on a weighted sum of the unsupervised loss and supervised loss) Luo in view of Kong in view of Huang and further in view of Zhu do not specifically disclose: generating a set of distorted labeled numerical embeddings based, at least in part, on the set of labeled numerical embeddings; generating, via the first transformer model, a fourth set of contextual numerical embeddings based, at least in part, on the set of distorted labeled numerical embeddings and the set of distorted unlabeled numerical embeddings; Li teaches: generating a set of distorted labeled [Li, 0049] discloses generating augmented (distorted) labelled batch and unlabeled batch to generate mixed augmented labelled batch X, and [0050] to calculate a total loss) generating, via the first transformer model, a fourth set of [Li, 0049] discloses generating augmented (distorted) labelled batch and unlabeled batch to generate mixed augmented labelled batch X, and [0050] to calculate a total loss. [0045] The combined batch is processed during the training to generate outputs) Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Luo, Kong, Huang, Zhu and Li to use the method of training a neural network using both distorted labeled data and distorted unlabeled data of Li to implement the tabular data machine learning method of the instant invention. The suggestion and/or motivation for doing so is to improve the robustness of the machine learning model and to improve generalization performance of the model [Li, 0026, 0054]. Regarding claim 7, Luo in view of Kong teaches: The computer-implemented method as claimed in claim 5, wherein the second set of semi-supervised learning operations comprises: initializing the second transformer model based, at least in part, on a set of second transformer weights; ([Kong, page 957, Algorithm 1, line 1-6] The parameters of the transformer encoder f and projection head g are initialized. The first and second neural network is taught by Luo) generating a set of labeled categorical embeddings based, at least in part, on the set of labeled categorical features; ([Kong, page 957, Algorithm 1, lines 18-26] Output embedding e i is generated using the encoder transformer f using the labeled dataset D l d . The first neural network processing categorical features and the second neural network processing numerical features is taught by Luo) generating a set of distorted unlabeled categorical features based, at least in part, on the set of unlabeled categorical features; ([Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating distorted feature by applying time warping augmentation and generating an output embedding using the encoder f θ ( x i ) ^   . The Transformer based encoder includes convolutional layers that generates intermediate embeddings when the signal passes through the convolutional layers) generating a set of distorted unlabeled categorical embeddings based, at least in part, on the set of distorted unlabeled categorical features; ([Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating distorted embedding e i ^ using the encoder f θ ( x i ) ^   . The Transformer based encoder includes convolutional layers that generates intermediate embeddings when the signal passes through the convolutional layers) generating, via the second transformer model, a third set of contextual categorical embeddings based, at least in part, on the set of labeled categorical embeddings and the set of unlabeled categorical embeddings; ([Kong, page 957, Algorithm 1, lines 18-26] Output embedding e i is generated using the encoder transformer f using the labeled dataset D l d . [Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating distorted embedding e i ^ using the encoder f θ ( x i ) ^   . The Transformer based encoder includes convolutional layers that generates intermediate embeddings when the signal passes through the convolutional layers) generating, via the second transformer model, a fourth set of contextual categorical embeddings based, at least in part,[Kong, page 957, Algorithm 1, lines 18-26] Output embedding e i is generated using the encoder transformer f using the labeled dataset D l d . [Kong, page 957, Algorithm 1, lines 11-12] and [page 955, left col, B. Transformer-Based Encoder, lines 1-27] discloses generating distorted embedding e i ^ using the encoder f θ ( x i ) ^   . The Transformer based encoder includes convolutional layers that generates intermediate embeddings when the signal passes through the convolutional layers) computing a semi-supervised loss by comparing the third set of contextual categorical embeddings and the fourth set of contextual categorical embeddings; and ([Kong, page 957, Algorithm 1, lines 18-26] and [page 953, right col, lines 1-10] Calculate the cross entropy loss function, and update networks f and g based on the semi-supervised (total) entropy loss function calculated based on a weighted sum of the unsupervised loss and supervised loss) fine-tuning the second transformer model based, at least in part, on the semi-supervised loss, wherein the fine-tuning comprises adjusting the set of second transformer weights. ([Kong, page 957, Algorithm 1, lines 18-26] and [page 953, right col, lines 1-10] Calculate the cross entropy loss function, and update networks f and g based on the semi-supervised (total) entropy loss function calculated based on a weighted sum of the unsupervised loss and supervised loss) Luo in view of Kong in view of Huang and further in view of Zhu do not specifically disclose: generating a set of distorted labeled categorical embeddings based, at least in part, on the set of labeled categorical embeddings; generating, via the second transformer model, a fourth set of contextual categorical embeddings based, at least in part, on the set of distorted labeled categorical embeddings and the set of distorted unlabeled categorical embeddings; Li teaches: generating a set of distorted labeled [Li, 0049] discloses generating augmented (distorted) labelled batch and unlabeled batch to generate mixed augmented labelled batch X, and [0050] to calculate a total loss) generating, via the second transformer model, a fourth set of contextual [Li, 0049] discloses generating augmented (distorted) labelled batch and unlabeled batch to generate mixed augmented labelled batch X, and [0050] to calculate a total loss. [0045] The combined batch is processed during the training to generate outputs) Claim 13 is a system claim which recites the same features as method claim 4 and rejected at least for the same reasons. Claim 16 is a system claim which recites the same features as method claim 7 and rejected at least for the same reasons. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUN KWON whose telephone number is (571)272-2072. The examiner can normally be reached Monday – Friday 8:00AM – 5:00PM ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Kawsar can be reached at (571)270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JUN KWON/Examiner, Art Unit 2127 /BRIAN M SMITH/Primary Examiner, Art Unit 2122
Read full office action

Prosecution Timeline

Jun 04, 2024
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737633
TRAINING A FEDERATED GENERATIVE ADVERSARIAL NETWORK
3y 9m to grant Granted Sep 15, 2026
Patent 12731035
LABEL INFERENCE IN SPLIT LEARNING DEFENSES
3y 8m to grant Granted Sep 08, 2026
Patent 12718063
DEEP LEARNING ARCHITECTURE FOR ADVERSE MEDIA SCREENING
3y 11m to grant Granted Aug 25, 2026
Patent 12711383
ACCURATE ENSEMBLE BY MUTATING NEURAL NETWORK PARAMETERS
7y 3m to grant Granted Aug 18, 2026
Patent 12705504
KNOWLEDGE BASE CONSTRUCTION
8y 5m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
41%
Grant Probability
88%
With Interview (+47.2%)
4y 8m (~2y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 78 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month