Prosecution Insights
Last updated: October 04, 2026
Application No. 18/176,380

Determining Machine Learning Model Performance on Unlabeled Out Of Distribution Data

Final Rejection §101§103
Filed
Feb 28, 2023
Examiner
MEYER, JACQUELINE CHRISTINE
Art Unit
2144
Tech Center
2100 — Computer Architecture & Software
Assignee
ORACLE INTERNATIONAL Corporation
OA Round
2 (Final)
65%
Grant Probability
Favorable
3-4
OA Rounds
4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 65% — above average
65%
Career Allowance Rate
15 granted / 23 resolved
+10.2% vs TC avg
Strong +62% interview lift
Without
With
+61.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
18 currently pending
Career history
42
Total Applications
across all art units

Statute-Specific Performance

§101
24.4%
-15.6% vs TC avg
§103
54.7%
+14.7% vs TC avg
§102
8.9%
-31.1% vs TC avg
§112
11.1%
-28.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 23 resolved cases

Office Action

§101 §103
Detailed Action Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea (mental process) without significantly more. Regarding claim 1, in Step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “A system, comprising: at least one processor; a memory…”. A system of the described configuration is within one of the four statutory categories of invention. In Step 2a Prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mental process but for recitation of generic computer components: “determine respective unbiased estimates for false positives, false negatives, true positives, and true negatives for performance of a machine learning model applied to a target dataset without corresponding ground truth labels according to importance sampling weights applied to the predictions of the machine learning model made given the items in the source dataset with respect to corresponding ground truth labels for the items in the source dataset” (This seems to be directed toward a mathematical process in view of the specification at [0014], [0016], & [0020-0032], and a mathematical process recites an abstract idea (MPEP 2106).) “determine, based on two or more of the respective unbiased estimates, a performance metric for the machine learning model on the target dataset without corresponding ground truth labels” (This seems to be directed toward a mathematical process in view of the specification at [0014], [0016], & [0020-0032], and a mathematical process recites an abstract idea (MPEP 2106).) If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. According, the claim “recites” an abstract idea. In Step 2a Prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: “A system, comprising: at least one processor; a memory, comprising program instructions that when executed by the at least one processor cause the at least one processor to implement a machine learning model evaluation system” (Uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).) “the machine learning model evaluation system configured to: receive, via an interface for the machine learning model evaluation system, a request for a performance metric” (Adding insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g)).) “obtain a source dataset with corresponding ground truth labels for items in the source dataset” (Adding insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g)).) “return, via the interface, the performance metric for the machine learning model on the target dataset” (Adding insignificant extra-solution activity (mere data output) to the judicial exception (MPEP 2106.05(g)).) Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In Step 2b of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional element (iii) recites use of a computer as a tool to perform the abstract idea, which is not indicative of significantly more. Additional elements (iv), (v), & (vi) recite insignificant extra-solution activities. Further, elements (iv) and (v) recite steps of receiving/transmitting data via a network, which has been determined by the courts to recite a well-understood, routine, and conventional activity, which is not indicative of significantly more (Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362). Further, element (vi) recites steps that present output of data which has been determined by the courts to recite a well-understood, routine, and conventional activity which is not indicative of significantly more (Presenting offers and gathering statistics, OIP Techs., 788 F.3d at 1362-63, 115 USPQ2d at 1092-93). Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Regarding claim 2, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 2 recites the following additional abstract ideas: “wherein the machine learning model evaluation system is further configured to: train a target density estimator for the target dataset to predict respective probabilities for items in the target dataset” (This is directed to a mathematical process in view of the specification at [0023-0030], and a mathematical process recites an abstract idea (MPEP 2106).) “train a source density estimator for the source dataset to predict respective probabilities for the items in the source dataset” (This is directed to a mathematical process in view of the specification at [0023-0030], and a mathematical process recites an abstract idea (MPEP 2106).) “determine the importance sampling weights applied to the predictions of the machine learning model using the target density estimator and the source density estimator.” (This is directed to a mathematical process in view of the specification at [0023-0030], and a mathematical process recites an abstract idea (MPEP 2106).) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 3, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 3 recites “wherein the target dataset is a simulated dataset” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 4, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 4 recites “wherein the performance metric for the machine learning model on the target dataset is a non-linear performance metric” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 5, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 5 recites “wherein the machine learning model is a document classifier” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 6, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 6 recites “wherein the performance metric for the machine learning model on the target dataset is a fairness metric” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 7, in Step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “A method”. A method is one of the four statutory categories of invention. In Step 2a Prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mental process but for recitation of generic computer components: “determining respective unbiased estimates for false positives, false negatives, true positives, and true negatives for performance of a machine learning model applied to a target dataset without corresponding ground truth labels according to importance sampling weights applied to the predictions of the machine learning model made given the items in the source dataset with respect to corresponding ground truth labels for the items in the source dataset” (This seems to be directed toward a mathematical process in view of the specification at [0014], [0016], & [0020-0032], and a mathematical process recites an abstract idea (MPEP 2106).) “determining, based on two or more of the respective unbiased estimates, a performance metric for the machine learning model on the target dataset without corresponding ground truth labels” (This seems to be directed toward a mathematical process in view of the specification at [0014], [0016], & [0020-0032], and a mathematical process recites an abstract idea (MPEP 2106).) If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. According, the claim “recites” an abstract idea. In Step 2a Prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: “performing, by one or more computing devices…” (Uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).) “obtaining a source dataset with corresponding ground truth labels for items in the source dataset” (Adding insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g)).) “providing, via an interface, the performance metric for the machine learning model on the target dataset” (Adding insignificant extra-solution activity (mere data output) to the judicial exception (MPEP 2106.05(g)).) Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In Step 2b of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional element (iii) recites use of a computer as a tool to perform the abstract idea, which is not indicative of significantly more. Additional elements (iv) & (v) recite insignificant extra-solution activities. Further, element (iv) recites steps of receiving/transmitting data via a network, which has been determined by the courts to recite a well-understood, routine, and conventional activity, which is not indicative of significantly more (Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362). Further, element (v) recites steps that present output of data which has been determined by the courts to recite a well-understood, routine, and conventional activity which is not indicative of significantly more (Presenting offers and gathering statistics, OIP Techs., 788 F.3d at 1362-63, 115 USPQ2d at 1092-93). Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Regarding claims 8-10, they are dependent upon claim 7, and thereby incorporate the limitations of, and corresponding analysis applied to claim 7. Further, claims 8-10 comprise similar additional limitations as claims 2, 4, & 6, respectively, and are rejected under the same rationale. Regarding claim 11, it is dependent upon claim 7, and thereby incorporates the limitations of, and corresponding analysis applied to claim 7. Further, claim 11 recites “wherein the target dataset is a synthetic dataset” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 12, it is dependent upon claim 7, and thereby incorporates the limitations of, and corresponding analysis applied to claim 7. Further, claim 12 recites the following additional abstract idea: “wherein the machine learning model performs named entity recognition” (This seems to be directed toward a mathematical process in view of the specification at [0014], [0016], & [0020-0033], and a mathematical process recites an abstract idea (MPEP 2106).) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 13, it is dependent upon claim 7, and thereby incorporates the limitations of, and corresponding analysis applied to claim 7. Further, claim 13 recites “receiving, via the interface, a request to provide the performance metric for the machine learning model on the target dataset” (In step2A, prong 2, this mere instructions to apply the judicial exception (MPEP 2106.05(f).) In step 2B, mere instructions to apply the judicial exception is not indicative of significantly more.) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 14, in Step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “One or more non-transitory, computer-readable storage media”. Non-transitory storage media is within one of the four statutory categories of invention. In Step 2a Prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mental process but for recitation of generic computer components: “determining respective unbiased estimates for false positives, false negatives, true positives, and true negatives for performance of a machine learning model applied to a target dataset without corresponding ground truth labels according to importance sampling weights applied to the predictions of the machine learning model made given the items in the source dataset with respect to corresponding ground truth labels for the items in the source dataset” (This seems to be directed toward a mathematical process in view of the specification at [0014], [0016], & [0020-0032], and a mathematical process recites an abstract idea (MPEP 2106).) “determining, based on two or more of the respective unbiased estimates, a performance metric for the machine learning model on the target dataset without corresponding ground truth labels” (This seems to be directed toward a mathematical process in view of the specification at [0014], [0016], & [0020-0032], and a mathematical process recites an abstract idea (MPEP 2106).) If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. According, the claim “recites” an abstract idea. In Step 2a Prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: “One or more non-transitory, computer-readable storage media, storing program instructions that when executed on or across one or more computing devices, cause the one or more computing devices to implement…” (Uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).) “obtaining a source dataset with corresponding ground truth labels for items in the source dataset” (Adding insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g)).) “providing, via an interface, the performance metric for the machine learning model on the target dataset” (Adding insignificant extra-solution activity (mere data output) to the judicial exception (MPEP 2106.05(g)).) Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In Step 2b of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional element (iii) recites use of a computer as a tool to perform the abstract idea, which is not indicative of significantly more. Additional elements (iv) & (v) recite insignificant extra-solution activities. Further, element (iv) recites steps of receiving/transmitting data via a network, which has been determined by the courts to recite a well-understood, routine, and conventional activity, which is not indicative of significantly more (Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362). Further, element (v) recites steps that present output of data which has been determined by the courts to recite a well-understood, routine, and conventional activity which is not indicative of significantly more (Presenting offers and gathering statistics, OIP Techs., 788 F.3d at 1362-63, 115 USPQ2d at 1092-93). Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Regarding claim 15, it is dependent upon claim 14, and thereby incorporates the limitations of, and corresponding analysis applied to claim 14. Further, claim 15 comprises similar additional limitations as claim 2, and is rejected under the same rationale. Regarding claim 16, it is dependent upon claim 14, and thereby incorporates the limitations of, and corresponding analysis applied to claim 14. Further, claim 16 recites “wherein the performance metric for the machine learning model on the target dataset is a linear performance metric” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claims 17-20, they are dependent upon claim 14, and thereby incorporate the limitations of, and corresponding analysis applied to claim 14. Further, claims 17-20 comprise similar additional limitations as claims 4, 11, 5 & 13, respectively, and are rejected under the same rationale. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1-2, 4, & 6, are rejected under 35 U.S.C. 103 as being unpatentable over Sugiyama, M. et al. “Covariate Shift Adaptation by Importance Weighted Cross Validation.” Available at https://dl.acm.org/doi/10.5555/1314498.1390324#:~:text=Under%20the%20covariate%20shift%2C%20standard,even%20under%20the%20covariate%20shift on 1 December 2007 (hereafter, SUGIYAMA), and further in view of Singh, H. “Everything you Need to Know About Hardware Requirements for Machine Learning.” Available at https://www.einfochips.com/blog/everything-you-need-to-know-about-hardware-requirements-for-machine-learning/ on 24 June 2019 (hereafter, SINGH), and further in view of Dudley, J. et al. “A Review of User Interface Design for Interactive Machine Learning.” Available at https://dl.acm.org/doi/10.1145/3185517 on 13 June 2018 (hereafter, DUDLEY) Regarding claim 1, SUGIYAMA teaches “receive… for the machine learning model evaluation system, a request for a performance metric”: ([Page 988, Equation 4] “ PNG media_image1.png 139 498 media_image1.png Greyscale In the above equation, xi and yi are labeled source examples. We see the Ptest(xi) and Ptrain(xi) being used in correlation as well. The equation is used to calculate an estimate of model performance (risk / expected loss) on an unlabeled target distribution, which is equivalent to receiving a request for a performance metric for a machine learning model evaluation system.”) Further, SUGIYAMA teaches “obtain a source dataset with corresponding ground truth labels for items in the source dataset”: ([Pages 986-987, 2.1 Supervised Learning under Covariate Shift] “…Let PNG media_image2.png 23 135 media_image2.png Greyscale be the training samples, where PNG media_image3.png 22 102 media_image3.png Greyscale is an i.i.d. training input point following a probability distribution Ptrain (x) and PNG media_image4.png 19 91 media_image4.png Greyscale is a corresponding training output value following a conditional probability distribution P(y|x). P(y|x) may be regarded as the sum of true output f(x) and noise.”) This citation along with the remainder of the cited segment explains mathematical processes where {xi} PNG media_image5.png 23 22 media_image5.png Greyscale represents the source dataset, the collection of input feature vectors drawn from the known distribution. In the context of the claim, the “source dataset” is simply the labeled datset used to compute importance-weighted estimates for the target. Further, each xi has an associated yi that is explicitly observed. These yi values are the ground-truth labels referenced in the claim. They are used in computing the weighted losses and he confusion estimates. In other words, the source dataset is “obtained” by having the training data available with observed labels, which correlates to the limitation as claimed. Further, SUGIYAMA teaches “determine respective unbiased estimates for false positives, false negatives, true positives, and true negatives for performance of a machine learning model applied to a target dataset without corresponding ground truth labels according to importance sampling weights applied to the predictions of the machine learning model made given the items in the source dataset with respect to corresponding ground truth labels for the items in the source dataset”: ([Pages 989-990, 3. Importance Weighted Cross Validation] “To compensate for the effect of the covariate shift in the CV procedure, we propose the following importance weighted CV (IWCV): PNG media_image6.png 171 592 media_image6.png Greyscale That is, the validation error in the CV procedure is weighted according to the importance. The following lemma shows that LOOIWCV gives an almost unbiased estimate of the risk even under the covariate shift… …the simple variant of LOOCV called LOOIWCV provides an unbiased estimate of the risk with n - 1 samples even under the covariate shift…”) The above cited text, along with the remainder of the cited segment of the reference teaches the limitation as claimed, especially when considered in view of the previous section 2 as well. f(xi) is a prediction of the model for source input xi. Ptest(xi) / Ptrain(xi) is an importance weight. PNG media_image7.png 17 11 media_image7.png Greyscale (f(xi),yi) is the loss function (0-1 loss for classification / mis-classification). 0-1 loss can be decomposed into true positive, false positive, true negative, false negative contributions by expressing the misclassification indicator as the sum of indicator functions over those four outcome types, because the importance-weighted expectation of any loss function, including the 0-1 loss, provides an unbiased estimate of the target-domain risk, as shown in the equations at [Pages 989-990, Section 3], and for each example, you can count whether it is a correct or incorrect prediction for each class. The importance weight ensures that these counts are unbiased with respect to the target distribution. By weighting the losses of the training examples according to the density ratio between target and source distributions, we can obtain an unbiased estimate of the expected loss of f on the target distribution. Further, SUGIYAMA teaches “determine, based on two or more of the respective unbiased estimates, a performance metric for the machine learning model on the target dataset without corresponding ground truth labels”: ([Pages 989-990, 3. Importance Weighted Cross Validation] “To compensate for the effect of the covariate shift in the CV procedure, we propose the following importance weighted CV (IWCV): PNG media_image6.png 171 592 media_image6.png Greyscale That is, the validation error in the CV procedure is weighted according to the importance. The following lemma shows that LOOIWCV gives an almost unbiased estimate of the risk even under the covariate shift… …the simple variant of LOOCV called LOOIWCV provides an unbiased estimate of the risk with n - 1 samples even under the covariate shift…”) The above cited text, along with the remainder of the cited segment of the reference teaches the limitation as claimed, especially when considered in view of the previous section 2 as well. f(xi) is a prediction of the model for source input xi. Ptest(xi) / Ptrain(xi) is an importance weight. PNG media_image7.png 17 11 media_image7.png Greyscale (f(xi),yi) is the loss function (0-1 loss for classification / mis-classification). 0-1 loss can be decomposed into true positive, false positive, true negative, false negative contributions, and for each example, you can count whether it is a correct or incorrect prediction for each class. The importance weight ensures that these counts are unbiased with respect to the target distribution. By weighting the losses of the training examples according to the density ratio between target and source distributions, we can obtain an unbiased estimate of the expected loss of f on the target distribution. Further, SUGIYAMA teaches “return, via the interface, the performance metric for the machine learning model on the target dataset”: ([Page 992, Figure 1] “ PNG media_image8.png 654 772 media_image8.png Greyscale In the above figure, we see the calculated performance metrics for the machine learning model on the target dataset, as previously cited, being displayed after being returned, by the previously cited functions, on an interface.”) SUGIYAMA fails to explicitly teach “A system, comprising: at least one processor; a memory, comprising program instructions that when executed by the at least one processor cause the at least one processor” and the request being received “…via an interface…”. However, analogous art, SINGH, does teach “A system, comprising: at least one processor; a memory, comprising program instructions that when executed by the at least one processor cause the at least one processor”: ([Hardware requirements for machine learning] “The first thing you should determine is what kind of resource does your task requires. Let’s have a look how different tasks will have different hardware requirements: If your tasks are small and can fit in a complex sequential processing, you don’t need a big system. You could even skip the GPUs altogether. A CPU such as i7–7500U can train an average of ~115 examples/second. So, if you are planning to work on other ML areas or algorithms, a GPU is not necessary. If your task is a bit intensive, and has a manageable data, a reasonably powerful GPU would be a better choice for you. A laptop with a dedicated graphics card of high end should do the work. There are a few high end (and expectedly heavy) laptops like Nvidia GTX 1080 (8 GB VRAM), which can train an average of ~14k examples/second. In addition, you can build your own PC with a reasonable CPU and a powerful GPU, but keep in mind that the CPU must not bottleneck the GPU. For instance, an i7-7500U will work flawlessly with a GTX 1080 GPU…”) This section describes a need for processing power of some degree (CPU or GPU) for machine learning tasks, giving multiple example systems, all with at least one processor. And further: ([Memory and Storage] “Machine learning models are getting larger, requiring more memory and storage capacity. High-bandwidth memory (HBM) and solid-state drives (SSDs) are becoming critical for training and running these models efficiently.”) This citation explicitly shows a requirement for the system to contain a memory. It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of SUGIYAMA with the teachings of SINGH because SUGIYAMA teaches of a machine learning model evaluation system methodology, while SINGH teaches of the necessary hardware requirements of using machine learning models. One of ordinary skill in the art would be motivated to do so because without meeting hardware requirements for any given computer application, the computer application will not function. SUGIYAMA in view of SINGH still fails to explicitly teach the request being received “…via an interface…”. However, analogous, art, DUDLEY, does teach this: ([Abstract] “Interactive Machine Learning (IML) seeks to complement human perception and intelligence by tightly integrating these strengths with the computational power and speed of computers. The interactive process is designed to involve input from the user (receiving a request from the user, via the interface) but does not require the background knowledge or experience that might be necessary to work with more traditional machine learning techniques… User interface design is fundamental to the success of this approach”) This citation shows that the purpose of DUDLEY is to investigate the use of user interfaces for machine learning processes, in order to make their use more accessible. It explicitly receives input from the user to activate and use the machine learning processes, which correlates to a request being received from an interface. It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of SUGIYAMA in view of SINGH with the teachings of DUDLEY because SUGIYAMA in view of SINGH teaches a system of determining and returning the performance of machine learning models via performance metrics, while DUDLEY teaches methods to streamline the use of machine learning processes using a user interface. One of ordinary skill in the art would be motivated to do so because, as DUDLEY points out in its abstract, “We propose and describe a structural and behavioural model of a generalised IML (Interactive Machine Learning) system and identify solution principles for building effective interfaces for IML.” Regarding claim 2, SUGIYAMA in view of SINGH & DUDLEY teaches the limitations of claim 1. Further, SUGIYAMA teaches “wherein the machine learning model evaluation system is further configured to: train a target density estimator for the target dataset to predict respective probabilities for items in the target dataset; train a source density estimator for the source dataset to predict respective probabilities for the items in the source dataset; and determine the importance sampling weights applied to the predictions of the machine learning model using the target density estimator and the source density estimator”: ([Pages 989-990, 3. Importance Weighted Cross Validation] “To compensate for the effect of the covariate shift in the CV procedure, we propose the following importance weighted CV (IWCV): PNG media_image6.png 171 592 media_image6.png Greyscale That is, the validation error in the CV procedure is weighted according to the importance. The following lemma shows that LOOIWCV gives an almost unbiased estimate of the risk even under the covariate shift… …the simple variant of LOOCV called LOOIWCV provides an unbiased estimate of the risk with n - 1 samples even under the covariate shift…”) The above cited text, along with the remainder of the cited segment of the reference teaches the limitation as claimed, especially when considered in view of the previous section 2 as well. f(xi) is a prediction of the model for source input xi. Ptest(xi) / Ptrain(xi) is an importance weight. PNG media_image7.png 17 11 media_image7.png Greyscale (f(xi),yi) is the loss function (0-1 loss for classification / mis-classification). 0-1 loss can be decomposed into true positive, false positive, true negative, false negative contributions by expressing the misclassification indicator as the sum of indicator functions over those four outcome types, because the importance-weighted expectation of any loss function, including the 0-1 loss, provides an unbiased estimate of the target-domain risk, as shown in the equations at [Pages 989-990, Section 3], and for each example, you can count whether it is a correct or incorrect prediction for each class. The importance weight ensures that these counts are unbiased with respect to the target distribution. By weighting the losses of the training examples according to the density ratio between target and source distributions, we can obtain an unbiased estimate of the expected loss of f on the target distribution. The cited “importance weights” correlate to the claimed “importance sampling weights”, which are applied to the predictions in the equations above using the estimators. And further: ([Page 992, Paragraph 2] “We also carried out the same simulation under unknown training and test densities (target and source density); they are estimated by maximum likelihood fitting of a single Gaussian model or a Gaussian kernel density estimator (simultaneously the target density estimator and the source density estimator) with variance determined by Silverman’s rule-of-thumb bandwidth selection rule... For estimating the test input density, we draw 100 unlabeled samples following Ptest(x). The simulation results had very similar trends to the case with known densities (therefore we omit the detail), although the error gets slightly larger. This implies that, for this toy regression problem, estimating the densities from data does not significantly degrade the quality of learning.”) The training and test densities each being estimated using the Gaussian kernel density estimator, according to Ptest(x) and Ptrain(x) shows the estimation of target and source density using a density estimator. Regarding claim 4, SUGIYAMA in view of SINGH & DUDLEY teaches the limitations of claim 1. Further, SUGIYAMA teaches “wherein the performance metric for the machine learning model on the target dataset is a non-linear performance metric”: ([Page 988, Equation 4] “ PNG media_image1.png 139 498 media_image1.png Greyscale In the above equation, xi and yi are labeled source examples. We see the Ptest(xi) and Ptrain(xi) being used in correlation as well. The equation is used to calculate an estimate of model performance (risk / expected loss) on an unlabeled target distribution, which is equivalent to receiving a request for a performance metric for a machine learning model evaluation system.”) And further: ([Page 986, Paragraph 3] “Model selection is one of the key ingredients in machine learning. However, under the covariate shift, a standard model selection technique such as cross validation (CV) (Stone, 1974; Wahba, 1990) does not work as desired; more specifically, the unbiasedness that guarantees the accuracy of CV does not hold under the covariate shift anymore. To cope with this problem, we propose a novel variant of CV called importance weighted CV (IWCV)… existing methods have a number of limitations, for example, in the loss function, parameter learning method, and model. In particular, the existing methods cannot be applied to classification scenarios. On the other hand, the proposed IWCV overcomes these limitations: it allows for any loss function, parameter learning method, and model;”) This citation shows that the performance metric taught by SUGIYAMA, based on a loss function, can be either non-linear or linear, based on the type of loss function/model used. The above citation, as used previously, shows the performance metric that is returned which is inherently dependent on whatever model is being used, and its loss function. This citations shows that the model being used could be any model with any loss function, which would encompass both non-linear and linear functions, resulting in the above cited performance metric being either non-linear or linear, just like the claimed invention. Regarding claim 6, SUGIYAMA in view of SINGH & DUDLEY teaches the limitations of claim 1. Further, SUGIYAMA teaches “wherein the performance metric for the machine learning model on the target dataset is a fairness metric”: ([Pages 989-990, 3. Importance Weighted Cross Validation] “To compensate for the effect of the covariate shift in the CV procedure, we propose the following importance weighted CV (IWCV): PNG media_image6.png 171 592 media_image6.png Greyscale That is, the validation error in the CV procedure is weighted according to the importance. The following lemma shows that LOOIWCV gives an almost unbiased estimate of the risk even under the covariate shift… …the simple variant of LOOCV called LOOIWCV provides an unbiased estimate of the risk with n - 1 samples even under the covariate shift…”) An unbiased estimate for the performance metric is a fair metric. The above cited text, along with the remainder of the cited segment of the reference teaches the limitation as claimed, especially when considered in view of the previous section 2 as well. f(xi) is a prediction of the model for source input xi. Ptest(xi) / Ptrain(xi) is an importance weight. PNG media_image7.png 17 11 media_image7.png Greyscale (f(xi),yi) is the loss function (0-1 loss for classification / mis-classification). 0-1 loss can be decomposed into true positive, false positive, true negative, false negative contributions, and for each example, you can count whether it is a correct or incorrect prediction for each class. The importance weight ensures that these counts are unbiased with respect to the target distribution. By weighting the losses of the training examples according to the density ratio between target and source distributions, we can obtain an unbiased estimate of the expected loss of f on the target distribution. The performance metric being used to weight it to be “unbiased” makes the performance metrics here “fair”. Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over SUGIYAMA in view of SINGH and DUDLEY, as applied to claims above, and further in view of Science Direct, “Simulated Data” Available online at https://www.sciencedirect.com/topics/computer-science/simulated-data on 30 November 2018 (hereafter, SCIENCEDIRECT). Regarding claim 3, SUGIYAMA in view of SINGH & DUDLEY teaches the limitations of claim 1. SUGIYAMA in view of SINGH & DUDLEY fails to explicitly teach “wherein the target dataset is a simulated dataset.” However, analogous art, SCIENCEDIRECT, does teach this: ([1. Introduction to Simulated Data in Computer Science] “Simulated data are microdata produced artificially from one or more statistical or computational models, rather than collected from real-world observations… Researchers frequently opt to create and use custom-generated simulated datasets tailored to the specific requirements of their experiments… where public datasets may not cover all scenarios. Simulated data offer advantages including ease of generation, limitless supply, pre-annotation, and reduced cost, and can circumvent ethical concerns and practical issues such as privacy and security. The availability of simulated data is especially important when real-world data are inaccessible due to confidentiality, privacy, or security issues, motivating the design of benchmark data for algorithm evaluation.”) And further: ([3. Applications of Simulated Data in Computer Science] “Simulated data are comparatively easier to generate, inexhaustible, pre-annotated, and less expensive than real data, while also enabling training on scenarios impractical or impossible to collect in the real world. Synthetic data and simulators catalyze richer representations and more sophisticated forms of learning, including multimodal, continual, and embodied learning, and are instrumental in the study of biological systems by enabling the training and testing of deep neural networks (DNNs). Progress in computer graphics tools and 3D assets has facilitated the synthesis of training data for tasks such as object detection, tracking, viewpoint estimation, semantic segmentation, robot manipulation, pose estimation, gaze estimation, and activity recognition, often improving DNN performance when combined with real data. Simulated data are also used to generate optimized stimuli for activating specific neural populations and can provide unique opportunities for direct comparisons of behavior and neural activation in artificial neural networks, humans, and nonhuman primates in simulated environments.”) It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of SUGIYAMA in view of SINGH and DUDLEY with the teachings of SCIENCEDIRECT because, as pointed out by SUGIYAMA in ([Page 986, Paragraph 3] the methods can be used with any model, and as pointed out by SCIENCEDIRECT, the use of simulated data with machine learning models is well-known in the art. One of ordinary skill in the art would be motivated to do so because, as pointed out by SCIENCEDIRECT above, there are numerous advantages to the use of synthetic data including “enabling training on scenarios impractical or impossible to collect in the real world”, “enabling the training and testing of deep neural networks (DNNs)”, and “often improving DNN performance when combined with real data. Simulated data are also used to generate optimized stimuli for activating specific neural populations and can provide unique opportunities for direct comparisons of behavior and neural activation in artificial neural networks” Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over SUGIYAMA in view of SINGH and DUDLEY, as applied to claims above, and further in view of Sealpath, “Advantages of Data Classification boosted by AI and Machine Learning” Available online at https://www.sealpath.com/blog/advantages-data-classification-ia-ml/ on 26 October 2022 (hereafter, SEALPATH). Regarding claim 5, SUGIYAMA in view of SINGH & DUDLEY teaches the limitations of claim 1. Further, SUGIYAMA in view of SINGH & DUDLEY fails to explicitly teach “wherein the machine learning model is a document classifier.” However, analogous art, SEALPATH, does teach this: ([IA and Machine Learning applied to Information Classification, Paragraphs 1-2] “Artificial Intelligence and machine learning techniques are of great value, helping to improve technology by providing security analysts with a faster and more efficient way of evaluating potential threats. Thanks to their application, they allow us to detect strange patterns or unusual user behaviour to help us anticipate attacks against our information. Thanks to machine learning algorithms, this advanced data classification technology can be applied, raising the accuracy level regarding the characteristics that make the specific contents of a file confidential. The models, taking advantage of machine learning, are previously trained for several years using data that contain personal information, medical, financial and other data. This previous learning helps predict the sensitivity level of previously non-labelled data.”) It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of SUGIYAMA in view of SINGH and DUDLEY with the teachings of SEALPATH because, as pointed out by SUGIYAMA in ([Page 986, Paragraph 3] the methods can be used with any model, and as pointed out by SEALPATH, the use of machine learning models as data classification models is well-known in the art. One of ordinary skill in the art would be motivated to do so because, as pointed out by SEALPATH, “this advanced data classification technology can be applied, raising the accuracy level regarding the characteristics that make the specific contents of a file confidential.” Claims 7-10, & 14-17 are rejected under 35 U.S.C. 103 as being unpatentable over SUGIYAMA, and further in view of SINGH. Regarding claim 7, SUGIYAMA teaches “obtaining a source dataset with corresponding ground truth labels for items in the source dataset”: ([Pages 986-987, 2.1 Supervised Learning under Covariate Shift] “…Let PNG media_image2.png 23 135 media_image2.png Greyscale be the training samples, where PNG media_image3.png 22 102 media_image3.png Greyscale is an i.i.d. training input point following a probability distribution Ptrain (x) and PNG media_image4.png 19 91 media_image4.png Greyscale is a corresponding training output value following a conditional probability distribution P(y|x). P(y|x) may be regarded as the sum of true output f(x) and noise.”) This citation along with the remainder of the cited segment explains mathematical processes where {xi} PNG media_image5.png 23 22 media_image5.png Greyscale represents the source dataset, the collection of input feature vectors drawn from the known distribution. In the context of the claim, the “source dataset” is simply the labeled datset used to compute importance-weighted estimates for the target. Further, each xi has an associated yi that is explicitly observed. These yi values are the ground-truth labels referenced in the claim. They are used in computing the weighted losses and he confusion estimates. In other words, the source dataset is “obtained” by having the training data available with observed labels, which correlates to the limitation as claimed. Further, SUGIYAMA teaches “determining respective unbiased estimates for false positives, false negatives, true positives, and true negatives for performance of a machine learning model applied to a target dataset without corresponding ground truth labels according to importance sampling weights applied to the predictions of the machine learning model made given the items in the source dataset with respect to corresponding ground truth labels for the items in the source dataset”: ([Pages 989-990, 3. Importance Weighted Cross Validation] “To compensate for the effect of the covariate shift in the CV procedure, we propose the following importance weighted CV (IWCV): PNG media_image6.png 171 592 media_image6.png Greyscale That is, the validation error in the CV procedure is weighted according to the importance. The following lemma shows that LOOIWCV gives an almost unbiased estimate of the risk even under the covariate shift… …the simple variant of LOOCV called LOOIWCV provides an unbiased estimate of the risk with n - 1 samples even under the covariate shift…”) The above cited text, along with the remainder of the cited segment of the reference teaches the limitation as claimed, especially when considered in view of the previous section 2 as well. f(xi) is a prediction of the model for source input xi. Ptest(xi) / Ptrain(xi) is an importance weight. PNG media_image7.png 17 11 media_image7.png Greyscale (f(xi),yi) is the loss function (0-1 loss for classification / mis-classification). 0-1 loss can be decomposed into true positive, false positive, true negative, false negative contributions by expressing the misclassification indicator as the sum of indicator functions over those four outcome types, because the importance-weighted expectation of any loss function, including the 0-1 loss, provides an unbiased estimate of the target-domain risk, as shown in the equations at [Pages 989-990, Section 3], and for each example, you can count whether it is a correct or incorrect prediction for each class. The importance weight ensures that these counts are unbiased with respect to the target distribution. By weighting the losses of the training examples according to the density ratio between target and source distributions, we can obtain an unbiased estimate of the expected loss of f on the target distribution. Further, SUGIYAMA teaches “determining, based on two or more of the respective unbiased estimates, a performance metric for the machine learning model on the target dataset without corresponding ground truth labels”: ([Pages 989-990, 3. Importance Weighted Cross Validation] “To compensate for the effect of the covariate shift in the CV procedure, we propose the following importance weighted CV (IWCV): PNG media_image6.png 171 592 media_image6.png Greyscale That is, the validation error in the CV procedure is weighted according to the importance. The following lemma shows that LOOIWCV gives an almost unbiased estimate of the risk even under the covariate shift… …the simple variant of LOOCV called LOOIWCV provides an unbiased estimate of the risk with n - 1 samples even under the covariate shift…”) The above cited text, along with the remainder of the cited segment of the reference teaches the limitation as claimed, especially when considered in view of the previous section 2 as well. f(xi) is a prediction of the model for source input xi. Ptest(xi) / Ptrain(xi) is an importance weight. PNG media_image7.png 17 11 media_image7.png Greyscale (f(xi),yi) is the loss function (0-1 loss for classification / mis-classification). 0-1 loss can be decomposed into true positive, false positive, true negative, false negative contributions, and for each example, you can count whether it is a correct or incorrect prediction for each class. The importance weight ensures that these counts are unbiased with respect to the target distribution. By weighting the losses of the training examples according to the density ratio between target and source distributions, we can obtain an unbiased estimate of the expected loss of f on the target distribution. Further, SUGIYAMA teaches “providing, via an interface, the performance metric for the machine learning model on the target dataset”: ([Page 992, Figure 1] “ PNG media_image8.png 654 772 media_image8.png Greyscale In the above figure, we see the calculated performance metrics for the machine learning model on the target dataset, as previously cited, being displayed after being returned, by the previously cited functions, on an interface.”) SUGIYAMA fails to explicitly teach “performing, by one or more computing devices:”. However, analogous art, SINGH, does teach this: ([Hardware requirements for machine learning] “The first thing you should determine is what kind of resource does your task requires. Let’s have a look how different tasks will have different hardware requirements: If your tasks are small and can fit in a complex sequential processing, you don’t need a big system. You could even skip the GPUs altogether. A CPU such as i7–7500U can train an average of ~115 examples/second. So, if you are planning to work on other ML areas or algorithms, a GPU is not necessary. If your task is a bit intensive, and has a manageable data, a reasonably powerful GPU would be a better choice for you. A laptop with a dedicated graphics card of high end should do the work. There are a few high end (and expectedly heavy) laptops like Nvidia GTX 1080 (8 GB VRAM), which can train an average of ~14k examples/second. In addition, you can build your own PC with a reasonable CPU and a powerful GPU, but keep in mind that the CPU must not bottleneck the GPU. For instance, an i7-7500U will work flawlessly with a GTX 1080 GPU…”) This section describes a need for processing power of some degree (CPU or GPU) for machine learning tasks, giving multiple example systems, all with at least one processor. It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of SUGIYAMA with the teachings of SINGH because SUGIYAMA teaches of a machine learning model evaluation system methodology, while SINGH teaches of the necessary hardware requirements of using machine learning models. One of ordinary skill in the art would be motivated to do so because without meeting hardware requirements for any given computer application, the computer application will not function. Regarding claim 8, SUGIYAMA in view of SINGH teaches the limitations of claim 7. Further, SUGIYAMA teaches “training a target density estimator for the target dataset to predict respective probabilities for items in the target dataset; training a source density estimator for the source dataset to predict respective probabilities for the items in the source dataset; and determining the importance sampling weights applied to the predictions of the machine learning model using the target density estimator and the source density estimator”: ([Pages 989-990, 3. Importance Weighted Cross Validation] “To compensate for the effect of the covariate shift in the CV procedure, we propose the following importance weighted CV (IWCV): PNG media_image6.png 171 592 media_image6.png Greyscale That is, the validation error in the CV procedure is weighted according to the importance. The following lemma shows that LOOIWCV gives an almost unbiased estimate of the risk even under the covariate shift… …the simple variant of LOOCV called LOOIWCV provides an unbiased estimate of the risk with n - 1 samples even under the covariate shift…”) The above cited text, along with the remainder of the cited segment of the reference teaches the limitation as claimed, especially when considered in view of the previous section 2 as well. f(xi) is a prediction of the model for source input xi. Ptest(xi) / Ptrain(xi) is an importance weight. PNG media_image7.png 17 11 media_image7.png Greyscale (f(xi),yi) is the loss function (0-1 loss for classification / mis-classification). 0-1 loss can be decomposed into true positive, false positive, true negative, false negative contributions by expressing the misclassification indicator as the sum of indicator functions over those four outcome types, because the importance-weighted expectation of any loss function, including the 0-1 loss, provides an unbiased estimate of the target-domain risk, as shown in the equations at [Pages 989-990, Section 3], and for each example, you can count whether it is a correct or incorrect prediction for each class. The importance weight ensures that these counts are unbiased with respect to the target distribution. By weighting the losses of the training examples according to the density ratio between target and source distributions, we can obtain an unbiased estimate of the expected loss of f on the target distribution. The cited “importance weights” correlate to the claimed “importance sampling weights”, which are applied to the predictions in the equations above using the estimators. And further: ([Page 992, Paragraph 2] “We also carried out the same simulation under unknown training and test densities (target and source density); they are estimated by maximum likelihood fitting of a single Gaussian model or a Gaussian kernel density estimator (simultaneously the target density estimator and the source density estimator) with variance determined by Silverman’s rule-of-thumb bandwidth selection rule... For estimating the test input density, we draw 100 unlabeled samples following Ptest(x). The simulation results had very similar trends to the case with known densities (therefore we omit the detail), although the error gets slightly larger. This implies that, for this toy regression problem, estimating the densities from data does not significantly degrade the quality of learning.”) The training and test densities each being estimated using the Gaussian kernel density estimator, according to Ptest(x) and Ptrain(x) shows the estimation of target and source density using a density estimator. Regarding claim 9, SUGIYAMA in view of SINGH teaches the limitations of claim 7. Further, SUGIYAMA teaches “wherein the performance metric for the machine learning model on the target dataset is a non-linear performance metric”: ([Page 988, Equation 4] “ PNG media_image1.png 139 498 media_image1.png Greyscale In the above equation, xi and yi are labeled source examples. We see the Ptest(xi) and Ptrain(xi) being used in correlation as well. The equation is used to calculate an estimate of model performance (risk / expected loss) on an unlabeled target distribution, which is equivalent to receiving a request for a performance metric for a machine learning model evaluation system.”) And further: ([Page 986, Paragraph 3] “Model selection is one of the key ingredients in machine learning. However, under the covariate shift, a standard model selection technique such as cross validation (CV) (Stone, 1974; Wahba, 1990) does not work as desired; more specifically, the unbiasedness that guarantees the accuracy of CV does not hold under the covariate shift anymore. To cope with this problem, we propose a novel variant of CV called importance weighted CV (IWCV)… existing methods have a number of limitations, for example, in the loss function, parameter learning method, and model. In particular, the existing methods cannot be applied to classification scenarios. On the other hand, the proposed IWCV overcomes these limitations: it allows for any loss function, parameter learning method, and model;”) This citation shows that the performance metric taught by SUGIYAMA, based on a loss function, can be either non-linear or linear, based on the type of loss function/model used. The above citation, as used previously, shows the performance metric that is returned which is inherently dependent on whatever model is being used, and its loss function. This citations shows that the model being used could be any model with any loss function, which would encompass both non-linear and linear functions, resulting in the above cited performance metric being either non-linear or linear, just like the claimed invention. Regarding claim 10, SUGIYAMA in view of SINGH teaches the limitations of claim 7. Further, SUGIYAMA teaches “wherein the performance metric for the machine learning model on the target dataset is a fairness metric”: ([Pages 989-990, 3. Importance Weighted Cross Validation] “To compensate for the effect of the covariate shift in the CV procedure, we propose the following importance weighted CV (IWCV): PNG media_image6.png 171 592 media_image6.png Greyscale That is, the validation error in the CV procedure is weighted according to the importance. The following lemma shows that LOOIWCV gives an almost unbiased estimate of the risk even under the covariate shift… …the simple variant of LOOCV called LOOIWCV provides an unbiased estimate of the risk with n - 1 samples even under the covariate shift…”) An unbiased estimate for the performance metric is a fair metric. The above cited text, along with the remainder of the cited segment of the reference teaches the limitation as claimed, especially when considered in view of the previous section 2 as well. f(xi) is a prediction of the model for source input xi. Ptest(xi) / Ptrain(xi) is an importance weight. PNG media_image7.png 17 11 media_image7.png Greyscale (f(xi),yi) is the loss function (0-1 loss for classification / mis-classification). 0-1 loss can be decomposed into true positive, false positive, true negative, false negative contributions, and for each example, you can count whether it is a correct or incorrect prediction for each class. The importance weight ensures that these counts are unbiased with respect to the target distribution. By weighting the losses of the training examples according to the density ratio between target and source distributions, we can obtain an unbiased estimate of the expected loss of f on the target distribution. The performance metric being used to weight it to be “unbiased” makes the performance metrics here “fair”. Regarding claim 14, SUGIYAMA teaches “obtaining a source dataset with corresponding ground truth labels for items in the source dataset”: ([Pages 986-987, 2.1 Supervised Learning under Covariate Shift] “…Let PNG media_image2.png 23 135 media_image2.png Greyscale be the training samples, where PNG media_image3.png 22 102 media_image3.png Greyscale is an i.i.d. training input point following a probability distribution Ptrain (x) and PNG media_image4.png 19 91 media_image4.png Greyscale is a corresponding training output value following a conditional probability distribution P(y|x). P(y|x) may be regarded as the sum of true output f(x) and noise.”) This citation along with the remainder of the cited segment explains mathematical processes where {xi} PNG media_image5.png 23 22 media_image5.png Greyscale represents the source dataset, the collection of input feature vectors drawn from the known distribution. In the context of the claim, the “source dataset” is simply the labeled datset used to compute importance-weighted estimates for the target. Further, each xi has an associated yi that is explicitly observed. These yi values are the ground-truth labels referenced in the claim. They are used in computing the weighted losses and he confusion estimates. In other words, the source dataset is “obtained” by having the training data available with observed labels, which correlates to the limitation as claimed. Further, SUGIYAMA teaches “determining respective unbiased estimates for false positives, false negatives, true positives, and true negatives for performance of a machine learning model applied to a target dataset without corresponding ground truth labels according to importance sampling weights applied to the predictions of the machine learning model made given the items in the source dataset with respect to corresponding ground truth labels for the items in the source dataset”: ([Pages 989-990, 3. Importance Weighted Cross Validation] “To compensate for the effect of the covariate shift in the CV procedure, we propose the following importance weighted CV (IWCV): PNG media_image6.png 171 592 media_image6.png Greyscale That is, the validation error in the CV procedure is weighted according to the importance. The following lemma shows that LOOIWCV gives an almost unbiased estimate of the risk even under the covariate shift… …the simple variant of LOOCV called LOOIWCV provides an unbiased estimate of the risk with n - 1 samples even under the covariate shift…”) The above cited text, along with the remainder of the cited segment of the reference teaches the limitation as claimed, especially when considered in view of the previous section 2 as well. f(xi) is a prediction of the model for source input xi. Ptest(xi) / Ptrain(xi) is an importance weight. PNG media_image7.png 17 11 media_image7.png Greyscale (f(xi),yi) is the loss function (0-1 loss for classification / mis-classification). 0-1 loss can be decomposed into true positive, false positive, true negative, false negative contributions by expressing the misclassification indicator as the sum of indicator functions over those four outcome types, because the importance-weighted expectation of any loss function, including the 0-1 loss, provides an unbiased estimate of the target-domain risk, as shown in the equations at [Pages 989-990, Section 3], and for each example, you can count whether it is a correct or incorrect prediction for each class. The importance weight ensures that these counts are unbiased with respect to the target distribution. By weighting the losses of the training examples according to the density ratio between target and source distributions, we can obtain an unbiased estimate of the expected loss of f on the target distribution. Further, SUGIYAMA teaches “determining, based on two or more of the respective unbiased estimates, a performance metric for the machine learning model on the target dataset without corresponding ground truth labels”: ([Pages 989-990, 3. Importance Weighted Cross Validation] “To compensate for the effect of the covariate shift in the CV procedure, we propose the following importance weighted CV (IWCV): PNG media_image6.png 171 592 media_image6.png Greyscale That is, the validation error in the CV procedure is weighted according to the importance. The following lemma shows that LOOIWCV gives an almost unbiased estimate of the risk even under the covariate shift… …the simple variant of LOOCV called LOOIWCV provides an unbiased estimate of the risk with n - 1 samples even under the covariate shift…”) The above cited text, along with the remainder of the cited segment of the reference teaches the limitation as claimed, especially when considered in view of the previous section 2 as well. f(xi) is a prediction of the model for source input xi. Ptest(xi) / Ptrain(xi) is an importance weight. PNG media_image7.png 17 11 media_image7.png Greyscale (f(xi),yi) is the loss function (0-1 loss for classification / mis-classification). 0-1 loss can be decomposed into true positive, false positive, true negative, false negative contributions, and for each example, you can count whether it is a correct or incorrect prediction for each class. The importance weight ensures that these counts are unbiased with respect to the target distribution. By weighting the losses of the training examples according to the density ratio between target and source distributions, we can obtain an unbiased estimate of the expected loss of f on the target distribution. Further, SUGIYAMA teaches “providing, via an interface, the performance metric for the machine learning model on the target dataset”: ([Page 992, Figure 1] “ PNG media_image8.png 654 772 media_image8.png Greyscale In the above figure, we see the calculated performance metrics for the machine learning model on the target dataset, as previously cited, being displayed after being returned, by the previously cited functions, on an interface.”) SUGIYAMA fails to explicitly teach “One or more non-transitory, computer-readable storage media, storing program instructions that when executed on or across one or more computing devices, cause the one or more computing devices to implement:”. However, analogous art, SINGH, does teach this: ([Hardware requirements for machine learning] “The first thing you should determine is what kind of resource does your task requires. Let’s have a look how different tasks will have different hardware requirements: If your tasks are small and can fit in a complex sequential processing, you don’t need a big system. You could even skip the GPUs altogether. A CPU such as i7–7500U can train an average of ~115 examples/second. So, if you are planning to work on other ML areas or algorithms, a GPU is not necessary. If your task is a bit intensive, and has a manageable data, a reasonably powerful GPU would be a better choice for you. A laptop (non-transitory system) with a dedicated graphics card of high end should do the work. There are a few high end (and expectedly heavy) laptops like Nvidia GTX 1080 (8 GB VRAM), which can train an average of ~14k examples/second. In addition, you can build your own PC with a reasonable CPU and a powerful GPU, but keep in mind that the CPU must not bottleneck the GPU. For instance, an i7-7500U will work flawlessly with a GTX 1080 GPU…”) This section describes a need for processing power of some degree (CPU or GPU) for machine learning tasks, giving multiple example systems, all with at least one processor. And further: ([Memory and Storage] “Machine learning models are getting larger, requiring more memory and storage capacity. High-bandwidth memory (HBM) and solid-state drives (SSDs) (non-transitory mediums) are becoming critical for training and running these models efficiently.”) This citation explicitly shows a requirement for the system to contain a memory, including non-transitory mediums. It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of SUGIYAMA with the teachings of SINGH because SUGIYAMA teaches of a machine learning model evaluation system methodology, while SINGH teaches of the necessary hardware requirements of using machine learning models. One of ordinary skill in the art would be motivated to do so because without meeting hardware requirements for any given computer application, the computer application will not function. Regarding claim 15, SUGIYAMA in view of SINGH teaches the limitations of claim 14. Further, claim 15 comprises similar additional limitations as claim 8, and is rejected under the same rationale. Regarding claim 16, SUGIYAMA in view of SINGH teaches the limitations of claim 14. Further, SUGIYAMA teaches “wherein the performance metric for the machine learning model on the target dataset is a linear performance metric”: ([Page 988, Equation 4] “ PNG media_image1.png 139 498 media_image1.png Greyscale In the above equation, xi and yi are labeled source examples. We see the Ptest(xi) and Ptrain(xi) being used in correlation as well. The equation is used to calculate an estimate of model performance (risk / expected loss) on an unlabeled target distribution, which is equivalent to receiving a request for a performance metric for a machine learning model evaluation system.”) And further: ([Page 986, Paragraph 3] “Model selection is one of the key ingredients in machine learning. However, under the covariate shift, a standard model selection technique such as cross validation (CV) (Stone, 1974; Wahba, 1990) does not work as desired; more specifically, the unbiasedness that guarantees the accuracy of CV does not hold under the covariate shift anymore. To cope with this problem, we propose a novel variant of CV called importance weighted CV (IWCV)… existing methods have a number of limitations, for example, in the loss function, parameter learning method, and model. In particular, the existing methods cannot be applied to classification scenarios. On the other hand, the proposed IWCV overcomes these limitations: it allows for any loss function, parameter learning method, and model;”) This citation shows that the performance metric taught by SUGIYAMA, based on a loss function, can be either non-linear or linear, based on the type of loss function/model used. The above citation, as used previously, shows the performance metric that is returned which is inherently dependent on whatever model is being used, and its loss function. This citations shows that the model being used could be any model with any loss function, which would encompass both non-linear and linear functions, resulting in the above cited performance metric being either non-linear or linear, just like the claimed invention. Regarding claim 17, SUGIYAMA in view of SINGH teaches the limitations of claim 14. Further, SUGIYAMA teaches “wherein the performance metric for the machine learning model on the target dataset is a non-linear performance metric”: ([Page 988, Equation 4] “ PNG media_image1.png 139 498 media_image1.png Greyscale In the above equation, xi and yi are labeled source examples. We see the Ptest(xi) and Ptrain(xi) being used in correlation as well. The equation is used to calculate an estimate of model performance (risk / expected loss) on an unlabeled target distribution, which is equivalent to receiving a request for a performance metric for a machine learning model evaluation system.”) And further: ([Page 986, Paragraph 3] “Model selection is one of the key ingredients in machine learning. However, under the covariate shift, a standard model selection technique such as cross validation (CV) (Stone, 1974; Wahba, 1990) does not work as desired; more specifically, the unbiasedness that guarantees the accuracy of CV does not hold under the covariate shift anymore. To cope with this problem, we propose a novel variant of CV called importance weighted CV (IWCV)… existing methods have a number of limitations, for example, in the loss function, parameter learning method, and model. In particular, the existing methods cannot be applied to classification scenarios. On the other hand, the proposed IWCV overcomes these limitations: it allows for any loss function, parameter learning method, and model;”) This citation shows that the performance metric taught by SUGIYAMA, based on a loss function, can be either non-linear or linear, based on the type of loss function/model used. The above citation, as used previously, shows the performance metric that is returned which is inherently dependent on whatever model is being used, and its loss function. This citations shows that the model being used could be any model with any loss function, which would encompass both non-linear and linear functions, resulting in the above cited performance metric being either non-linear or linear, just like the claimed invention. Claims 11 & 18 are rejected under 35 U.S.C. 103 as being unpatentable over SUGIYAMA in view of SINGH, as applied to claims above, and further in view of SCIENCEDIRECT. Regarding claim 11, SUGIYAMA in view of SINGH teaches the limitations of claim 7. SUGIYAMA in view of SINGH fails to explicitly teach “wherein the target dataset is a synthetic dataset.” However, analogous art, SCIENCEDIRECT, does teach this: ([1. Introduction to Simulated Data in Computer Science] “Simulated data are microdata produced artificially from one or more statistical or computational models, rather than collected from real-world observations… Researchers frequently opt to create and use custom-generated simulated datasets tailored to the specific requirements of their experiments… where public datasets may not cover all scenarios. Simulated data offer advantages including ease of generation, limitless supply, pre-annotation, and reduced cost, and can circumvent ethical concerns and practical issues such as privacy and security. The availability of simulated data is especially important when real-world data are inaccessible due to confidentiality, privacy, or security issues, motivating the design of benchmark data for algorithm evaluation.”) And further: ([3. Applications of Simulated Data in Computer Science] “Simulated data are comparatively easier to generate, inexhaustible, pre-annotated, and less expensive than real data, while also enabling training on scenarios impractical or impossible to collect in the real world. Synthetic data and simulators catalyze richer representations and more sophisticated forms of learning, including multimodal, continual, and embodied learning, and are instrumental in the study of biological systems by enabling the training and testing of deep neural networks (DNNs). Progress in computer graphics tools and 3D assets has facilitated the synthesis of training data for tasks such as object detection, tracking, viewpoint estimation, semantic segmentation, robot manipulation, pose estimation, gaze estimation, and activity recognition, often improving DNN performance when combined with real data. Simulated data are also used to generate optimized stimuli for activating specific neural populations and can provide unique opportunities for direct comparisons of behavior and neural activation in artificial neural networks, humans, and nonhuman primates in simulated environments.”) Synthetic data is a type of simulated data, and the terms are often used interchangeably, as can be seen in the citations above. It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of SUGIYAMA in view of SINGH with the teachings of SCIENCEDIRECT because, as pointed out by SUGIYAMA in ([Page 986, Paragraph 3] the methods can be used with any model, and as pointed out by SCIENCEDIRECT, the use of simulated data with machine learning models is well-known in the art. One of ordinary skill in the art would be motivated to do so because, as pointed out by SCIENCEDIRECT above, there are numerous advantages to the use of synthetic data including “enabling training on scenarios impractical or impossible to collect in the real world”, “enabling the training and testing of deep neural networks (DNNs)”, and “often improving DNN performance when combined with real data. Simulated data are also used to generate optimized stimuli for activating specific neural populations and can provide unique opportunities for direct comparisons of behavior and neural activation in artificial neural networks” Regarding claim 18, SUGIYAMA in view of SINGH teaches the limitations of claim 14. Further, claim 18 comprises similar additional limitations as claim 11, and is rejected under the same rationale. Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over SUGIYAMA in view of SINGH, as applied to claims above, and further in view of Wu, Y, et al. “Clinical Named Entity Recognition Using Deep Learning Models.” Available online at https://pmc.ncbi.nlm.nih.gov/articles/PMC5977567/ on 16 April 2018 (hereafter, WU). Regarding claim 12, SUGIYAMA in view of SINGH teaches the limitations of claim 7. SUGIYAMA in view of SINGH fails to explicitly teach “wherein the machine learning model performs named entity recognition.” However, analogous art, WU, does teach this: ([Abstract] “Clinical Named Entity Recognition (NER) is a critical natural language processing (NLP) task to extract important concepts (named entities) from clinical narratives. Researchers have extensively investigated machine learning models for clinical NER. Recently, there have been increasing efforts to apply deep learning models to improve the performance of current clinical NER systems. This study examined two popular deep learning architectures, the Convolutional Neural Network (CNN) and the Recurrent Neural Network (RNN), to extract concepts from clinical texts. We compared the two deep neural network architectures with three baseline Conditional Random Fields (CRFs) models and two state-of-the-art clinical NER systems using the i2b2 2010 clinical concept extraction corpus. The evaluation results showed that the RNN model trained with the word embeddings achieved a new state-of-the- art performance (a strict F1 score of 85.94%) for the defined clinical NER task, outperforming the best-reported system that used both manually defined and unsupervised learning features. This study demonstrates the advantage of using deep neural network architectures for clinical concept extraction, including distributed feature representation, automatic feature learning, and long-term dependencies capture. This is one of the first studies to compare the two widely used deep learning models and demonstrate the superior performance of the RNN model for clinical NER.”) This citation shows that named entity recognition is a known use of machine learning models. It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of SUGIYAMA in view of SINGH with the teachings of WU because, as pointed out by SUGIYAMA in ([Page 986, Paragraph 3] the methods can be used with any model, and as pointed out by WU, the use of named entity recognition with machine learning models is well-known in the art. One of ordinary skill in the art would be motivated to do so because, as pointed out by WU in the abstract above, “The evaluation results showed that the RNN model trained with the word embeddings achieved a new state-of-the- art performance (a strict F1 score of 85.94%) for the defined clinical NER task, outperforming the best-reported system that used both manually defined and unsupervised learning features. This study demonstrates the advantage of using deep neural network architectures for clinical concept extraction, including distributed feature representation, automatic feature learning, and long-term dependencies capture.” Claims 13 & 20 are rejected under 35 U.S.C. 103 as being unpatentable over SUGIYAMA in view of SINGH, as applied to claims above, and further in view of DUDLEY. Regarding claim 13, SUGIYAMA in view of SINGH teaches the limitations of claim 7. Further, SUGIYAMA teaches “receiving… a request to provide the performance metric for the machine learning model on the target dataset”: ([Page 988, Equation 4] “ PNG media_image1.png 139 498 media_image1.png Greyscale In the above equation, xi and yi are labeled source examples. We see the Ptest(xi) and Ptrain(xi) being used in correlation as well. The equation is used to calculate an estimate of model performance (risk / expected loss) on an unlabeled target distribution, which is equivalent to receiving a request for a performance metric for a machine learning model evaluation system on a target dataset.”) SUGIYAMA in view of SINGH fails to explicitly teach the request being received “…via the interface…” However, analogous art, DUDLEY, does teach this: ([Abstract] “Interactive Machine Learning (IML) seeks to complement human perception and intelligence by tightly integrating these strengths with the computational power and speed of computers. The interactive process is designed to involve input from the user (receiving a request from the user, via the interface) but does not require the background knowledge or experience that might be necessary to work with more traditional machine learning techniques… User interface design is fundamental to the success of this approach”) This citation shows that the purpose of DUDLEY is to investigate the use of user interfaces for machine learning processes, in order to make their use more accessible. It explicitly receives input from the user to activate and use the machine learning processes, which correlates to a request being received from an interface. It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of SUGIYAMA in view of SINGH with the teachings of DUDLEY because SUGIYAMA in view of SINGH teaches a system of determining and returning the performance of machine learning models via performance metrics, while DUDLEY teaches methods to streamline the use of machine learning processes using a user interface. One of ordinary skill in the art would be motivated to do so because, as DUDLEY points out in its abstract, “We propose and describe a structural and behavioural model of a generalised IML (Interactive Machine Learning) system and identify solution principles for building effective interfaces for IML.” Regarding claim 20, SUGIYAMA in view of SINGH teaches the limitations of claim 14. Further, claim 20 comprises similar additional limitations as claim 13, and is rejected under the same rationale. Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over SUGIYAMA in view of SINGH, as applied to claims above, and further in view of SEALPATH. Regarding claim 19, SUGIYAMA in view of SINGH teaches the limitations of claim 14. Further, SUGIYAMA in view of SINGH fails to explicitly teach “wherein the machine learning model is a document classifier.” However, analogous art, SEALPATH, does teach this: ([IA and Machine Learning applied to Information Classification, Paragraphs 1-2] “Artificial Intelligence and machine learning techniques are of great value, helping to improve technology by providing security analysts with a faster and more efficient way of evaluating potential threats. Thanks to their application, they allow us to detect strange patterns or unusual user behaviour to help us anticipate attacks against our information. Thanks to machine learning algorithms, this advanced data classification technology can be applied, raising the accuracy level regarding the characteristics that make the specific contents of a file confidential. The models, taking advantage of machine learning, are previously trained for several years using data that contain personal information, medical, financial and other data. This previous learning helps predict the sensitivity level of previously non-labelled data.”) It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of SUGIYAMA in view of SINGH with the teachings of SEALPATH because, as pointed out by SUGIYAMA in ([Page 986, Paragraph 3] the methods can be used with any model, and as pointed out by SEALPATH, the use of machine learning models as data classification models is well-known in the art. One of ordinary skill in the art would be motivated to do so because, as pointed out by SEALPATH, “this advanced data classification technology can be applied, raising the accuracy level regarding the characteristics that make the specific contents of a file confidential.” Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW LEE LEWIS whose telephone number is (571)272-1906. The examiner can normally be reached Monday: 9:30AM - 3:30PM and Tuesday - Friday: 9:30AM - 6PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at (571)272-4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Matthew Lee Lewis/Examiner, Art Unit 2144 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2144
Read full office action

Prosecution Timeline

Feb 28, 2023
Application Filed
Oct 23, 2025
Non-Final Rejection mailed — §101, §103
Jan 24, 2026
Response Filed
Oct 01, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743602
Automated Variational Inference using Stochastic Models with Irregular Beliefs
3y 8m to grant Granted Sep 22, 2026
Patent 12737637
MACHINE LEARNING MODEL TRAINING WITH ADVERSARIAL LEARNING AND TRIPLET LOSS REGULARIZATION
3y 6m to grant Granted Sep 15, 2026
Patent 12717933
METHOD AND SYSTEM FOR SECURING NEURAL NETWORK MODELS
4y 2m to grant Granted Aug 25, 2026
Patent 12718094
SYSTEM AND METHOD FOR CONTINUAL REFINABLE NETWORK
3y 8m to grant Granted Aug 25, 2026
Patent 12705498
SYSTEMS AND METHODS FOR FEDERATED VALIDATION OF MODELS
3y 8m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
65%
Grant Probability
99%
With Interview (+61.7%)
3y 11m (~4m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 23 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month