Prosecution Insights
Last updated: October 02, 2026
Application No. 18/453,854

SYSTEMS AND METHODS FOR DEPLOYING MACHINE LEARNING MODELS TRAINED ON SYNTHETIC DATA GENERATED BASED ON A PREDICTED FUTURE DATA DRIFT

Final Rejection §101§103
Filed
Aug 22, 2023
Examiner
VO, STEVEN
Art Unit
2148
Tech Center
2100 — Computer Architecture & Software
Assignee
Capital One Services LLC
OA Round
2 (Final)
Grant Probability
Favorable
3-4
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
15 currently pending
Career history
10
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103
DETAILED ACTION This action is in response to the application files on 08/22/2023. Claims 1-20 are pending and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 3, 5, 7-12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gama et al. (A Survey on Concept Drift Adaptation) (hereafter referred to as Gama) in view of Li et al. (DDG-DA: Data Distribution Generation for Predictable Concept Drift Adaptation) (hereafter referred to as Li), Xu et al. (Modeling Tabular Data Using Conditional GAN) (hereafter referred to as Xu), and Dwork et al. (Calibrating Noise to Sensitivity in Private Data Analysis) (hereafter referred to as Dwork). Regarding claim 3, Gama teaches determining, based on a first data profile for a first snapshot of a data stream captured at a first time and a second data profile for a second snapshot of the data stream captured at a second time, that a first data drift between the first snapshot and the second snapshot exceeds a threshold (Gama, page 44:17, section 3.2.3, “These methods typically use a fixed reference window that summarizes the past information and a sliding detection window over the most recent examples” and “The ADaptive sliding WINdow or ADWIN [Bifet and Gavalda 2006, 2007] is another change detector using a detection window... Whenever two large enough (sub)windows of W exhibit distinct enough means, the algorithm concludes that the expected values within those windows are different” (Gama, page 44:18, section 3.2.3). Examiner notes that the windows are being mapped to first and second snapshot and ADWIN is being mapped to the threshold of the first data drift). in response to determining that the first data drift exceeds the threshold, updating a machine learning model based on the second snapshot to generate a first updated machine learning model, wherein the machine learning model was previously trained on the first snapshot (Gama, page 44:18, section 3.2.3, “Large enough and distinct enough are defined by the Hoeffding bound, testing whether the average of the two (sub)windows is larger than ε-cut” and “Retraining approaches need some data buffer to be stored in memory. Retraining has been used to emulate incremental learning with batch-learning algorithms [Gama et al. 2004]. At the beginning, a model is trained with all the available data. Next, whenever new data arrives, the previous model is discarded, the new data is merged with the previous data, and a new model is learned on this data” (Gama, page 44:20, section 3.3.1). Examiner notes that the Hoeffding bound is being mapped to how the first data drift is determined). determining, based on the second data profile and a third data profile for a third snapshot of the data stream captured at the third time, that a second data drift between the second snapshot and the third snapshot exceeds the threshold (Gama, page 44:17, section 3.2.3, “These methods typically use a fixed reference window that summarizes the past information and a sliding detection window over the most recent examples” and “The ADaptive sliding WINdow or ADWIN [Bifet and Gavalda 2006, 2007] is another change detector using a detection window... Whenever two large enough (sub)windows of W exhibit distinct enough means, the algorithm concludes that the expected values within those windows are different” (Gama, page 44:18, section 3.2.3). Examiner notes that the windows are being mapped to first and second snapshot and ADWIN is being mapped to the threshold of the first data drift). in response to determining that the second data drift exceeds the threshold, deploying the speculative machine learning model to replace the first updated machine learning model (Gama, page 44:18, section 3.2.3, “Large enough and distinct enough are defined by the Hoeffding bound, testing whether the average of the two (sub)windows is larger than ε-cut” and “Retraining approaches need some data buffer to be stored in memory. Retraining has been used to emulate incremental learning with batch-learning algorithms [Gama et al. 2004]. At the beginning, a model is trained with all the available data. Next, whenever new data arrives, the previous model is discarded, the new data is merged with the previous data, and a new model is learned on this data” (Gamma, page 44:20, section 3.3.1). Examiner notes that the Hoeffding bound is being mapped to how the first data drift is determined. Gama does not teach, but Li does teach extrapolating the first data drift to determine a predicted data drift corresponding to a third time in the future and generating, based on the predicted data drift, a synthetic snapshot of the data stream corresponding to the third time in the future, by one or more of replacing a feature, in the second snapshot of the data stream, with synthetic data, iteratively sampling from the first snapshot of the data stream and the second snapshot of the data stream, or adding random noise, to the synthetic snapshot, based on an amount of random noise required based on the first data profile and the second data profile (Li, section 1, “In practice, DDG-DA is designed as a dynamic data generator that can create sample data from previously observed data by following predicted future data distribution. In other words, DDG-DA generates the resampling probability of each historical data sample to construct the future data distribution in estimation”, “DDG-DA is expected to guide the model learning process in each task (t) by forecasting test data distribution. Historical data distribution information is useful to predict the target distribution of D(t) test and is input into DDG-DA. DDG-DA will learn concept drift patterns from training tasks and help to adapt models in test tasks” (Li, section 3.2), and “DDG-DA generates the resampling probability of each historical data sample to construct the future data distribution in estimation … the data comes in a sequential mode, meaning that, at any timestamp t, the learning task can only observe the new coming information, i.e., Data(t), in addition to historical ones, i.e., Data(t−1), Data(t−2), ···, Data(1)” (Li, section 1). generating a speculative machine learning model from updating the first updated machine learning model based on the synthetic snapshot (Li, section 4.1, “DDG-DA solves such a problem and performs best. It models the trend of concept drifts and generates a new dataset whose distribution is closer to that in the future. Then the forecasting model is trained on the new dataset to handle the future concept drift”). Gama and Li are considered analogous to the claimed invention because they both deal with data drift. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama to use the Data Distribution Generation for Predicable Concept Drift Adaptation (DDG-DA) from Li. Li teaches that “Only a few adaptation methods are model agnostic. DDG-DA focuses on predicting future data distribution, which is a model-agnostic solution and can benefit different customized forecasting models on streaming data.” (Li, section 2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Gama and Li do not teach, but Xu does teach extrapolating the first data drift to determine a predicted data drift corresponding to a third time in the future and generating, based on the predicted data drift, a synthetic snapshot of the data stream corresponding to the third time in the future, by one or more of replacing a feature, in the second snapshot of the data stream, with synthetic data, iteratively sampling from the first snapshot of the data stream and the second snapshot of the data stream, or adding random noise, to the synthetic snapshot, based on an amount of random noise required based on the first data profile and the second data profile (Xu, page 2, section 3, “After training [data synthesizer] G on [training set on table T] Ttrain, [synthetic table] Tsyn is constructed by independently sampling rows using G” and (Xu, page 4, section 4.3) PNG media_image1.png 254 564 media_image1.png Greyscale Examiner notes that the discrete is being mapped to the feature). Gama, Li, and Xu are considered analogous to the claimed invention because they deal with synthetic data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama and Li to include Xu’s conditional tabular generative adversarial network (CTGAN) method. Xu teaches that “CTGAN outperforms Bayesian methods on most of the real datasets whereas other deep learning methods could not” (Xu, page 1, Abstract) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Gama, Li, and Xu do not teach, but Dwork does teach extrapolating the first data drift to determine a predicted data drift corresponding to a third time in the future and generating, based on the predicted data drift, a synthetic snapshot of the data stream corresponding to the third time in the future, by one or more of replacing a feature, in the second snapshot of the data stream, with synthetic data, iteratively sampling from the first snapshot of the data stream and the second snapshot of the data stream, or adding random noise, to the synthetic snapshot, based on an amount of random noise required based on the first data profile and the second data profile (Dwork, page 265, section 1.1, “we analyze the sensitivity of specific data analysis functions, including histograms, contingency tables, and covariance matrices, all of which have very high-dimensional output, and show that their sensitivities are independent of the dimension. Previous privacy-preserving approximations to these quantities used noise proportional to the dimension; the new analysis permits noise of size O(1)” and “the true answer is perturbed by the addition of random noise generated according to a carefully chosen distribution, and this response, the true answer plus noise, is returned to the user” (Dwork, page 265, section 1)). Gama, Li, Xu, and Dwork are considered analogous to the claimed invention because they deal with synthetic data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama and Li to include Dwork’s method of calibrating the standard deviation of noise based on sensitivity. Dwork teaches that “The new analysis shows that for several particular applications substantially less noise is needed than was previously understood to be the case.” (Dwork, page 265, Abstract) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Regarding claim 5, Gama, Li, Xu, and Dwork teach the method of claim 3, Gama further teaches removing the first data profile after a time period, wherein the time period relates to the time needed to replace the first updated machine learning model (Gama, page 44:18, section 3.2.3, “Whenever two large enough (sub)windows of W exhibit distinct enough means, the algorithm concludes that the expected values within those windows are different, and the older (sub)window is dropped”). Regarding claim 7, Gama, Li, Xu, and Dwork teach the method of claim 3, Gama further teaches the threshold is determined based on size or frequency of a data profile, and wherein determining the first data drift exceeds the threshold is based on a similarity between the first data profile and the second data profile (Gama, page 44:18, section, “Whenever two large enough (sub)windows of W exhibit distinct enough means, the algorithm concludes that the expected values within those windows are different, and the older (sub)window is dropped. Large enough and distinct enough are defined by the Hoeffding bound, testing whether the average of the two (sub)windows is larger than εcut”. Examiner notes that large enough is mapped to the size of the data while distinct enough is mapped to the similarity between the data). Regarding claim 8, Gama, Li, Xu, and Dwork teach the method of claim 3, Gama further teaches determining a performance threshold, wherein the performance threshold is related to a performance metric of a machine learning model (Gama, page 44:16, section 3.2.2, “A commonly used confidence level for Warning is 95% with the threshold pi + σi ≥ pmin + 2 ∗ σmin, and for Out-of-Control is 99% with the threshold pi +σi ≥ pmin +3∗σmin”. Examiner notes that pi +σi ≥ pmin +3∗σmin is mapped to the performance threshold where pi +σi is that performance metric and pmin +3∗σmin is that threshold”). in response to determining the machine learning model is not above the performance threshold, updating the machine learning model based on a new snapshot from the data stream (Gama, page 44:20, section 3.3.1, “Retraining has been used to emulate incremental learning with batch-learning algorithms [Gama et al. 2004]. At the beginning, a model is trained with all the available data. Next, whenever new data arrives, the previous model is discarded, the new data is merged with the previous data, and a new model is learned on this data”). Regarding claim 9, Gama, Li, Xu, and Dwork teach the method of claim 3, Gama further teaches determining, based on the first data profile and the second data profile, that a first data drift between the first snapshot and the second snapshot exceeds a threshold further comprises detecting a deviation in a distribution of attributes between the first data profile and the second data profile (Gama, page 44:4, section 2.1, “Formally, concept drift between time point t0 and time point t1 can be defined as ∃X: pt0 (X, y) != pt1 (X, y) where pt0 denotes the joint distribution at time t0 between the set of input variables X and the target variable y. Changes in data can be characterized as changes in the components of this relation”. Examiner notes that pt0 (X, y) != pt1 (X, y) shows that there is a deviation between the two data profiles). Regarding claim 10, Gama, Li, Xu, and Dwork teach the method of claim 3, Gama, Li, and Dwork do not teach, but Xu teaches generating, based on the predicted data drift, a synthetic snapshot of the data stream corresponding to the third time in the future further comprises: identifying a feature to replace in the second snapshot based on the predicted data drift (Xu, page 4, section 4.3, PNG media_image1.png 254 564 media_image1.png Greyscale Examiner notes that the discrete is being mapped to the feature). generating synthetic data to replace the feature in the second snapshot, wherein the synthetic data has a distribution of values that are similar to the second snapshot and are able to generate the predicted data drift (Xu, page 4, section 4.3, “Integrating a conditional generator into the architecture of a GAN requires to deal with the following issues: 1) it is necessary to devise a representation for the condition as well as to prepare an input for it, 2) it is necessary for the generated rows to preserve the condition as it is given, and 3) it is necessary for the conditional generator to learn the real data conditional distribution, i.e. PG(row|Di∗ = k∗) = P(row|Di∗ = k∗), so that we can reconstruct the original distribution as P(row) = Σ PG(row|Di∗ = k∗)P(Di∗ = k)”. Examiner notes that step 2 shows that the rows generated must be preserved or similar to the original data). replacing the feature in the second snapshot with the synthetic data (Xu, page 2, section 3, “After training [data synthesizer] G on [training set on table T] Ttrain, [synthetic table] Tsyn is constructed by independently sampling rows using G”). Gama, Li, and Xu are considered analogous to the claimed invention because they deal with synthetic data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama and Li to include Xu’s conditional tabular generative adversarial network (CTGAN) method. Xu teaches that “CTGAN outperforms Bayesian methods on most of the real datasets whereas other deep learning methods could not” (Xu, page 1, Abstract) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Regarding claim 11, Gama, Li, Xu, and Dwork teach the methods of claim 10, Gama, Xu, and Dwork do not teach, but Li teaches iteratively sampling from the first snapshot and the second snapshot to generate synthetic data for the predicted data drift (Li, section 1, “DDG-DA generates the resampling probability of each historical data sample to construct the future data distribution in estimation … the data comes in a sequential mode, meaning that, at any timestamp t, the learning task can only observe the new coming information, i.e., Data(t), in addition to historical ones, i.e., Data(t−1), Data(t−2), ···, Data(1)”). combining the synthetic data to generate a synthetic snapshot (Li, section 1, “DDG-DA generates the resampling probability of each historical data sample to construct the future data distribution in estimation”). Gama, Li, and Xu are considered analogous to the claimed invention because they all deal with synthetic data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama and Li to include Li’s DDG-DA method to iterate through a dataset. Li teaches that “Only a few adaptation methods are model agnostic. DDG-DA focuses on predicting future data distribution, which is a model-agnostic solution and can benefit different customized forecasting models on streaming data.” (Li, section 2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Regarding claim 12, Gama, Li, Xu, and Dwork teach the methods of claim 3, Gama, Li, and Xu do not teach, but Dwork teaches generating, based on the predicted data drift, a synthetic snapshot of the data stream corresponding to the third time in the future further comprises: processing the first data profile and the second data profile to determine an amount of random noise required, wherein the amount of random noise maintains the privacy of a dataset while minimizing the amount of variation of the dataset (Dwork, page 265, section 1.1, “we analyze the sensitivity of specific data analysis functions, including histograms, contingency tables, and covariance matrices, all of which have very high-dimensional output, and show that their sensitivities are independent of the dimension. Previous privacy-preserving approximations to these quantities used noise proportional to the dimension; the new analysis permits noise of size O(1)”) generating the random noise, wherein the random noise comprises additional data (Dwork, page 265, section 1, “We assume the database is held by a trusted server. On input a query function f mapping databases to reals, the so-called true answer is the result of applying f to the database. To protect privacy, the true answer is perturbed by the addition of random noise generated according to a carefully chosen distribution, and this response, the true answer plus noise, is returned to the user”). adding the random noise to the dataset by modifying data in the dataset to the additional data (Dwork, page 265, section 1, “the true answer is perturbed by the addition of random noise generated according to a carefully chosen distribution, and this response, the true answer plus noise, is returned to the user”). generating the synthetic snapshot based on the dataset (Dwork, page 267, section 1.1, “In the non-interactive set ting, the data collector—a trusted entity—publishes a “sanitized” version of the collected data; the literature uses terms such as “anonymization” and “de-identification”. Traditionally, sanitization employed some perturbation and data modification techniques”. Examiner notes that the sanitized version of the data is mapped to the synthetic snapshot). Gama, Li, and Xu are considered analogous to the claimed invention because they deal with synthetic data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama and Li to include Dwork’s method of calibrating the standard deviation of noise based on sensitivity. Dwork teaches that “The new analysis shows that for several particular applications substantially less noise is needed than was previously understood to be the case.” (Dwork, page 265, Abstract) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Claim(s) 1, 2, 4, 6, 13-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gama in view of Li, Xu, and Dwork, and Deosthali et al. (US 20240086762 A1) (hereafter referred to as Deosthali). Regarding claim 1, Gama teaches receiving a first snapshot of a data stream captured at a first time and a second snapshot of the data stream captured at a second time (Gama, page 44:4, section 2.1, “In the setting that we are considering, data arrives online, often in real time, forming a stream that is potentially infinite” and “Formally, concept drift between time point t0 and time point t1 can be defined as ∃X: pt0 (X, y) != pt1 (X, y)”. Examiner notes that the time points t0 and t1 are mapped to the first and second snapshot, respectively). processing, using a data profiler function, the first snapshot of the data stream, to generate a first data profile for the first snapshot, and the second snapshot of the data stream, to generate a second data profile for the second snapshot (Gama, page 44:17, bottom of the page, “The detection windows are used for estimating data distribution and are typically paired” and “Formally, concept drift between time point t0 and time point t1 can be defined as ∃X: pt0 (X, y) != pt1 (X, y), (2) where pt0 denotes the joint distribution at time t0 between the set of input variables X and the target variable y” (Gama, page 44:4, section 2.1). Examiner notes that the detection windows are processed to generate data distributions, which are mapped to data profiles). determining, based on the first data profile and the second data profile, that a first data drift between the first snapshot and the second snapshot exceeds a threshold (Gama, page 40:18, section 3.2.3, “Whenever two large enough (sub)windows of W exhibit distinct enough means, the algorithm concludes that the expected values within those windows are different, and the older (sub)window is dropped. Large enough and distinct enough are defined by the Hoeffding bound, testing whether the average of the two (sub)windows is larger than ε-cut”. Examiner notes that the Hoeffding bound is being mapped to how that data drift threshold was reached) in response to determining that the first data drift exceeds the threshold, updating a machine learning model based on the second snapshot to generate a first updated machine learning model, wherein the machine learning model was previously trained on the first snapshot (Gama, page 44:18, section 3.2.3, “Large enough and distinct enough are defined by the Hoeffding bound, testing whether the average of the two (sub)windows is larger than ε-cut” and “Retraining approaches need some data buffer to be stored in memory. Retraining has been used to emulate incremental learning with batch-learning algorithms [Gama et al. 2004]. At the beginning, a model is trained with all the available data. Next, whenever new data arrives, the previous model is discarded, the new data is merged with the previous data, and a new model is learned on this data” (Gama, page 44:20, section 3.3.1). Examiner notes that the Hoeffding bound is being mapped to how the first data drift is determined). receiving a third snapshot of the data stream captured at the third time (Gama, page 44:4, section 2.1, “In the setting that we are considering, data arrives online, often in real time, forming a stream that is potentially infinite”. Examiner notes that forming an infinite stream means that there can be multiple snapshots). processing, using the data profiler function, the third snapshot of the data stream to generate a third data profile for the third snapshot (Gama, page 44:17, bottom of the page, “The detection windows are used for estimating data distribution and are typically paired” and “Formally, concept drift between time point t0 and time point t1 can be defined as ∃X: pt0 (X, y) != pt1 (X, y), (2) where pt0 denotes the joint distribution at time t0 between the set of input variables X and the target variable y” (Gama, page 44:4, section 2.1). Examiner notes that the detection windows are processed to generate data distributions, which are mapped to data profiles). determining, based on the second data profile and a third data profile for a third snapshot of the data stream captured at the third time, that a second data drift between the second snapshot and the third snapshot exceeds the threshold (Gama, page 44:17, section 3.2.3, “These methods typically use a fixed reference window that summarizes the past information and a sliding detection window over the most recent examples” and “The ADaptive sliding WINdow or ADWIN [Bifet and Gavalda 2006, 2007] is another change detector using a detection window... Whenever two large enough (sub)windows of W exhibit distinct enough means, the algorithm concludes that the expected values within those windows are different” (Gama, page 44:18, section 3.2.3). Examiner notes that the windows are being mapped to first and second snapshot and ADWIN is being mapped to the threshold of the first data drift). in response to determining that the second data drift exceeds the threshold, deploying the speculative machine learning model to replace the first updated machine learning model (Gama, page 44:18, section 3.2.3, “Large enough and distinct enough are defined by the Hoeffding bound, testing whether the average of the two (sub)windows is larger than ε-cut” and “Retraining approaches need some data buffer to be stored in memory. Retraining has been used to emulate incremental learning with batch-learning algorithms [Gama et al. 2004]. At the beginning, a model is trained with all the available data. Next, whenever new data arrives, the previous model is discarded, the new data is merged with the previous data, and a new model is learned on this data” (Gamma, page 44:20, section 3.3.1). Examiner notes that the Hoeffding bound is being mapped to how the first data drift is determined. Gama does not teach, but Li does teach extrapolating the first data drift to determine a predicted data drift corresponding to a third time in the future and generating, based on the predicted data drift, a synthetic snapshot of the data stream corresponding to the third time in the future, by one or more of replacing a feature, in the second snapshot of the data stream, with synthetic data, iteratively sampling from the first snapshot of the data stream and the second snapshot of the data stream, or adding random noise, to the synthetic snapshot, based on an amount of random noise required based on the first data profile and the second data profile (Li, section 1, “In practice, DDG-DA is designed as a dynamic data generator that can create sample data from previously observed data by following predicted future data distribution. In other words, DDG-DA generates the resampling probability of each historical data sample to construct the future data distribution in estimation”, “DDG-DA is expected to guide the model learning process in each task (t) by forecasting test data distribution. Historical data distribution information is useful to predict the target distribution of D(t) test and is input into DDG-DA. DDG-DA will learn concept drift patterns from training tasks and help to adapt models in test tasks” (Li, section 3.2), and “DDG-DA generates the resampling probability of each historical data sample to construct the future data distribution in estimation … the data comes in a sequential mode, meaning that, at any timestamp t, the learning task can only observe the new coming information, i.e., Data(t), in addition to historical ones, i.e., Data(t−1), Data(t−2), ···, Data(1)” (Li, section 1). generating a speculative machine learning model from updating the first updated machine learning model based on the synthetic snapshot (Li, section 4.1, “DDG-DA solves such a problem and performs best. It models the trend of concept drifts and generates a new dataset whose distribution is closer to that in the future. Then the forecasting model is trained on the new dataset to handle the future concept drift”). Gama and Li are considered analogous to the claimed invention because they both deal with data drift. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama to use the Data Distribution Generation for Predicable Concept Drift Adaptation (DDG-DA) from Li. Li teaches that “Only a few adaptation methods are model agnostic. DDG-DA focuses on predicting future data distribution, which is a model-agnostic solution and can benefit different customized forecasting models on streaming data.” (Li, section 2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Gama and Li do not teach, but Xu does teach extrapolating the first data drift to determine a predicted data drift corresponding to a third time in the future and generating, based on the predicted data drift, a synthetic snapshot of the data stream corresponding to the third time in the future, by one or more of replacing a feature, in the second snapshot of the data stream, with synthetic data, iteratively sampling from the first snapshot of the data stream and the second snapshot of the data stream, or adding random noise, to the synthetic snapshot, based on an amount of random noise required based on the first data profile and the second data profile (Xu, page 2, section 3, “After training [data synthesizer] G on [training set on table T] Ttrain, [synthetic table] Tsyn is constructed by independently sampling rows using G” and (Xu, page 4, section 4.3) PNG media_image1.png 254 564 media_image1.png Greyscale Examiner notes that the discrete is being mapped to the feature). Gama, Li, and Xu are considered analogous to the claimed invention because they deal with synthetic data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama and Li to include Xu’s conditional tabular generative adversarial network (CTGAN) method. Xu teaches that “CTGAN outperforms Bayesian methods on most of the real datasets whereas other deep learning methods could not” (Xu, page 1, Abstract) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Gama, Li, and Xu do not teach, but Dwork does teach extrapolating the first data drift to determine a predicted data drift corresponding to a third time in the future and generating, based on the predicted data drift, a synthetic snapshot of the data stream corresponding to the third time in the future, by one or more of replacing a feature, in the second snapshot of the data stream, with synthetic data, iteratively sampling from the first snapshot of the data stream and the second snapshot of the data stream, or adding random noise, to the synthetic snapshot, based on an amount of random noise required based on the first data profile and the second data profile (Dwork, page 265, section 1.1, “we analyze the sensitivity of specific data analysis functions, including histograms, contingency tables, and covariance matrices, all of which have very high-dimensional output, and show that their sensitivities are independent of the dimension. Previous privacy-preserving approximations to these quantities used noise proportional to the dimension; the new analysis permits noise of size O(1)” and “the true answer is perturbed by the addition of random noise generated according to a carefully chosen distribution, and this response, the true answer plus noise, is returned to the user” (Dwork, page 265, section 1)). Gama, Li, Xu, and Dwork are considered analogous to the claimed invention because they deal with synthetic data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama and Li to include Dwork’s method of calibrating the standard deviation of noise based on sensitivity. Dwork teaches that “The new analysis shows that for several particular applications substantially less noise is needed than was previously understood to be the case.” (Dwork, page 265, Abstract) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Gama, Li, Xu, and Dwork do not teach, but Deosthali does teach one or more processors (Deosthali, page 2, paragraph 0030, “FIG. 2 illustrates a machine learning system 105 in accordance with one or more embodiments. The machine learning system 105 can be the same or similar to that previously described above. Embodiments of the machine learning system 105 can be implemented on one or more digital devices. A digital device can be any hardware device that includes a processor”). one or more non-transitory, computer-readable media storing instructions that, when executed by the one or more processors, cause operations (Deosthali, page 9, paragraph 0084, “In an embodiment, a non-transitory computer readable storage medium comprises instructions which, when executed by one or more hardware processors, causes performance of any of the operations described herein and/or recited in any of the claims”.) Gama, Li, Xu, Dwork, and Deosthali are considered analogous to the claimed invention because they deal with data drift. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama, Li, Xu, and Dwork to be able to run on a processor and non-transitory computer-readable medium. One of the ordinary skill in the art would have known to apply the known technique of using processors and a non-transitory computer-readable medium to perform Gama and Li’s techniques. Therefore, applying Li’s technique would yield the predictable result of running instructions on processors and a non-transitory computer-readable medium (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results. Regarding claim 2, Gama, Li, Xu, Dwork, and Deosthali teach the system of claim 1, Gama further teaches in response to determining that the second data drift exceeds the threshold, updating the first updated machine learning model based on the third snapshot to generate a second updated machine learning model (Gama, page 44:20, section 3.3.1, “Retraining approaches need some data buffer to be stored in memory. Retraining has been used to emulate incremental learning with batch-learning algorithms [Gama et al. 2004]. At the beginning, a model is trained with all the available data. Next, whenever new data arrives, the previous model is discarded, the new data is merged with the previous data, and a new model is learned on this data”). Gama does not teach, but Deosthali does teach subsequent to the second updated machine learning model being generated: determining whether a performance metric of the second updated machine learning model exceeds a performance metric of the speculative machine learning model (Deosthali, page 6, paragraph 0058, “different models can be compared (e.g., A/B-tested) based on their respective accuracy and performance to identify the best model across different model attributes and data distributions”) in response to determining that the performance metric of the second updated machine learning model does not exceed the performance metric of the speculative machine learning model, relabeling the speculative machine learning model to be the second updated machine learning model (Desothali, page 6, paragraph 0059, “At block 360, the process 300 selects a machine learning for deployment. User can save the best model to create the baseline for further data and model experiments. Saving the base model saves the model attributes, hyperattributes, runtime information, and the model file for scoring new datasets. A most optimized model that accommodates changes in data with robust model metrics across wide ranges of data and with satisfactory non-functional (latency, cost) performance metrics can be chosen for deployment”. Examiner notes that “saving” the model is mapped to relabeling the model). in response to determining that the performance metric of the second updated machine learning model exceeds the performance metric of the speculative machine learning model, deploying the second updated machine learning model to replace the speculative machine learning model (Deosthali, page 2, paragraph 0027, “Based on the evaluation information, the machine learning system 105 can analyze the effect of drift and select best-performing machine learning model 115 for deployment to the production system 107 from the baseline model or the candidate models”). Gama, Li, Xu, Dwork, and Deosthali are considered analogous to the claimed invention because they deal with data drift. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama, Li, Xu, and Dwork to determine which machine learning model to use from Deosthali. One of the ordinary skill in the art would have known to apply the known technique of comparing machine learning models based on accuracy and performance. Therefore, applying Li’s technique would yield the predictable result of determining the best performing machine learning model to use (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results. Regarding claim 4: Claim 4 recites the following contingent limitations: b - “in response to determining that the performance metric of the second updated machine learning model does not exceed the performance metric of the speculative machine learning model, relabeling the speculative machine learning model to be the second updated machine learning model” and c - “in response to determining that the performance metric of the second updated machine learning model exceeds the performance metric of the speculative machine learning model, deploying the second updated machine learning model to replace the speculative machine learning model”. These limitations are contingent because they recite steps that are only required to be performed in their conditions precedent are met. Limitation b only needs to be performed if the performance metric does not exceed a threshold and limitation c only needs to be performed if the performance metric does exceed a threshold. Therefore, the BRI under claim 4 requires only one of either limitation b or c. The examiner will use limitation c. Gama, Li, Xu, and Dwork teach the method of claim 3, Gama further teaches in response to determining that the second data drift exceeds the threshold, updating the first updated machine learning model based on the third snapshot to generate a second updated machine learning model (Gama, page 44:20, section 3.3.1, “Retraining approaches need some data buffer to be stored in memory. Retraining has been used to emulate incremental learning with batch-learning algorithms [Gama et al. 2004]. At the beginning, a model is trained with all the available data. Next, whenever new data arrives, the previous model is discarded, the new data is merged with the previous data, and a new model is learned on this data”). Gama does not teach, but Deosthali does teach subsequent to the second updated machine learning model being generated: determining whether a performance metric of the second updated machine learning model exceeds a performance metric of the speculative machine learning model (Deosthali, page 6, paragraph 0058, “different models can be compared (e.g., A/B-tested) based on their respective accuracy and performance to identify the best model across different model attributes and data distributions”) in response to determining that the performance metric of the second updated machine learning model does not exceed the performance metric of the speculative machine learning model, relabeling the speculative machine learning model to be the second updated machine learning model in response to determining that the performance metric of the second updated machine learning model exceeds the performance metric of the speculative machine learning model, deploying the second updated machine learning model to replace the speculative machine learning model (Deosthali, page 2, paragraph 0027, “Based on the evaluation information, the machine learning system 105 can analyze the effect of drift and select best-performing machine learning model 115 for deployment to the production system 107 from the baseline model or the candidate models”). Gama, Li, Xu, Dwork, and Deosthali are considered analogous to the claimed invention because they deal with data drift. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama, Li, Xu, and Dwork to determine which machine learning model to use from Deosthali. One of the ordinary skill in the art would have known to apply the known technique of comparing machine learning models based on accuracy and performance. Therefore, applying Li’s technique would yield the predictable result of determining the best performing machine learning model to use (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results. Regarding claim 6, Gama, Li, Xu, and Dwork teach the method of claim 3, Gama further teaches comparing the first data profile to the second data profile to process the first data drift using a detection algorithm, wherein the first data profile comprises a first plurality of metrics, wherein the first plurality of metrics comprises statistics of the first data profile, wherein the second data profile comprises a second plurality of metrics, and wherein the second plurality of metrics comprises statistics of the second data profile (Gama, page 44:18, section 3.2.3, “Whenever two large enough (sub)windows of W exhibit distinct enough means, the algorithm concludes that the expected values within those windows are different, and the older (sub)window is dropped. Large enough and distinct enough are defined by the Hoeffding bound, testing whether the average of the two (sub)windows is larger than ε-cut”. Examiner notes that the Hoeffding bound is the detecting algorithm). Gama does not teach, but Deosthali does teach comparing the first data profile to the second data profile to process the first data drift using a detection algorithm, wherein the first data profile comprises a first plurality of metrics, wherein the first plurality of metrics comprises statistics of the first data profile, wherein the second data profile comprises a second plurality of metrics, and wherein the second plurality of metrics comprises statistics of the second data profile (Deosthali, page 7, paragraph 0064, “the data validator 227 can generate a list of metrics of divergent dataset 215 attributes, such as distribution, mean, median mode, standard deviation, quantities of values in the dataset, and missing values, observations within one standard deviations (%), observations within two standard deviations (%), observations within three standard deviations (%), negative outliers (%), positive outliers (%), missing values (%), minimum value, maximum value, negative value, unique values, and zero values”). Gama, Li, Xu, Dwork, and Deosthali are considered analogous to the claimed invention because they deal with data drift. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama, Li, Xu, and Dwork to include a plurality of metrics from Deosthali to compare the sub windows from Gama. One of the ordinary skill in the art would have known to apply the known technique of using multiple metrics to compare datasets. Therefore, applying Li’s technique would yield the predictable result of accurately comparing datasets based on multiple variables (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results. Gama, Xu, and Dwork do not teach, but Li does teach determining the predicted data drift using a prediction machine learning model, wherein the prediction machine learning model is trained using a similarity between the first plurality of metrics and the second plurality of metrics, and wherein the prediction machine learning model is validated using the similarity between the second plurality of metrics and a third plurality of metrics, wherein the third plurality of metrics comprises statistics of the third data profile (Li, Abstract, “We first train a predictor to estimate the future data distribution, then leverage it to generate training samples, and finally train models on the generated data” and ” DDG-DA generates the resampling probability of each historical data sample to construct the future data distribution in estimation. However, it is quite challenging in reality to train this data generator to maximize the similarity between the predicted data distribution (represented by weighted resampling on historical data) and the ground truth future data distribution (represented by data in the future)” (Li, Abstract)). Gama, Li, Xu, Dwork, and Deosthali are considered analogous to the claimed invention because they deal with data drift. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama, Xu, Dwork, and Deosthali to include Li’s prediction machine learning model. One of the ordinary skill in the art would have known to apply the known technique of predicting data based on past data. Therefore, applying Li’s technique would yield the predictable result of accurately predict future data (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results. Regarding claim 13, Gama teaches determining, based on a first data profile for a first snapshot of a data stream captured at a first time and a second data profile for a second snapshot of the data stream captured at a second time, that a first data drift between the first snapshot and the second snapshot exceeds a threshold (Gama, page 44:17, section 3.2.3, “These methods typically use a fixed reference window that summarizes the past information and a sliding detection window over the most recent examples” and “The ADaptive sliding WINdow or ADWIN [Bifet and Gavalda 2006, 2007] is another change detector using a detection window... Whenever two large enough (sub)windows of W exhibit distinct enough means, the algorithm concludes that the expected values within those windows are different” (Gama, page 44:18, section 3.2.3). Examiner notes that the windows are being mapped to first and second snapshot and ADWIN is being mapped to the threshold of the first data drift). determining, based on the second data profile and a third data profile for a third snapshot of the data stream captured at the third time, that a second data drift between the second snapshot and the third snapshot exceeds the threshold (Gama, page 44:17, section 3.2.3, “These methods typically use a fixed reference window that summarizes the past information and a sliding detection window over the most recent examples” and “The ADaptive sliding WINdow or ADWIN [Bifet and Gavalda 2006, 2007] is another change detector using a detection window... Whenever two large enough (sub)windows of W exhibit distinct enough means, the algorithm concludes that the expected values within those windows are different” (Gama, page 44:18, section 3.2.3). Examiner notes that the windows are being mapped to first and second snapshot and ADWIN is being mapped to the threshold of the first data drift). in response to determining that the second data drift exceeds the threshold, deploying the speculative machine learning model Gama, page 44:20, section 3.3.1, “Retraining approaches need some data buffer to be stored in memory. Retraining has been used to emulate incremental learning with batch-learning algorithms [Gama et al. 2004]. At the beginning, a model is trained with all the available data. Next, whenever new data arrives, the previous model is discarded, the new data is merged with the previous data, and a new model is learned on this data”). Gama does not teach, but Li does teach extrapolating the first data drift to determine a predicted data drift corresponding to a third time in the future and generating, based on the predicted data drift, a synthetic snapshot of the data stream corresponding to the third time in the future, by one or more of replacing a feature, in the second snapshot of the data stream, with synthetic data, iteratively sampling from the first snapshot of the data stream and the second snapshot of the data stream, or adding random noise, to the synthetic snapshot, based on an amount of random noise required based on the first data profile and the second data profile (Li, section 1, “In practice, DDG-DA is designed as a dynamic data generator that can create sample data from previously observed data by following predicted future data distribution. In other words, DDG-DA generates the resampling probability of each historical data sample to construct the future data distribution in estimation”, “DDG-DA is expected to guide the model learning process in each task (t) by forecasting test data distribution. Historical data distribution information is useful to predict the target distribution of D(t) test and is input into DDG-DA. DDG-DA will learn concept drift patterns from training tasks and help to adapt models in test tasks” (Li, section 3.2), and “DDG-DA generates the resampling probability of each historical data sample to construct the future data distribution in estimation … the data comes in a sequential mode, meaning that, at any timestamp t, the learning task can only observe the new coming information, i.e., Data(t), in addition to historical ones, i.e., Data(t−1), Data(t−2), ···, Data(1)” (Li, section 1). generating a speculative machine learning model from updating the first updated machine learning model based on the synthetic snapshot (Li, section 4.1, “DDG-DA solves such a problem and performs best. It models the trend of concept drifts and generates a new dataset whose distribution is closer to that in the future. Then the forecasting model is trained on the new dataset to handle the future concept drift”). Gama and Li are considered analogous to the claimed invention because they both deal with data drift. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama to use the Data Distribution Generation for Predicable Concept Drift Adaptation (DDG-DA) from Li. Li teaches that “Only a few adaptation methods are model agnostic. DDG-DA focuses on predicting future data distribution, which is a model-agnostic solution and can benefit different customized forecasting models on streaming data.” (Li, section 2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Gama and Li do not teach, but Xu does teach extrapolating the first data drift to determine a predicted data drift corresponding to a third time in the future and generating, based on the predicted data drift, a synthetic snapshot of the data stream corresponding to the third time in the future, by one or more of replacing a feature, in the second snapshot of the data stream, with synthetic data, iteratively sampling from the first snapshot of the data stream and the second snapshot of the data stream, or adding random noise, to the synthetic snapshot, based on an amount of random noise required based on the first data profile and the second data profile (Xu, page 2, section 3, “After training [data synthesizer] G on [training set on table T] Ttrain, [synthetic table] Tsyn is constructed by independently sampling rows using G” and (Xu, page 4, section 4.3) PNG media_image1.png 254 564 media_image1.png Greyscale Examiner notes that the discrete is being mapped to the feature). Gama, Li, and Xu are considered analogous to the claimed invention because they deal with synthetic data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama and Li to include Xu’s conditional tabular generative adversarial network (CTGAN) method. Xu teaches that “CTGAN outperforms Bayesian methods on most of the real datasets whereas other deep learning methods could not” (Xu, page 1, Abstract) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Gama, Li, and Xu do not teach, but Dwork does teach extrapolating the first data drift to determine a predicted data drift corresponding to a third time in the future and generating, based on the predicted data drift, a synthetic snapshot of the data stream corresponding to the third time in the future, by one or more of replacing a feature, in the second snapshot of the data stream, with synthetic data, iteratively sampling from the first snapshot of the data stream and the second snapshot of the data stream, or adding random noise, to the synthetic snapshot, based on an amount of random noise required based on the first data profile and the second data profile (Dwork, page 265, section 1.1, “we analyze the sensitivity of specific data analysis functions, including histograms, contingency tables, and covariance matrices, all of which have very high-dimensional output, and show that their sensitivities are independent of the dimension. Previous privacy-preserving approximations to these quantities used noise proportional to the dimension; the new analysis permits noise of size O(1)” and “the true answer is perturbed by the addition of random noise generated according to a carefully chosen distribution, and this response, the true answer plus noise, is returned to the user” (Dwork, page 265, section 1)). Gama, Li, Xu, and Dwork are considered analogous to the claimed invention because they deal with synthetic data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama and Li to include Dwork’s method of calibrating the standard deviation of noise based on sensitivity. Dwork teaches that “The new analysis shows that for several particular applications substantially less noise is needed than was previously understood to be the case.” (Dwork, page 265, Abstract) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Gama, Li, Xu, and Dwork do not teach, but Deosthali does teach One or more non-transitory, computer-readable storage media storing instructions that, when executed by one or more processors, cause operations (Deosthali, page 9, paragraph 0084, “In an embodiment, a non-transitory computer readable storage medium comprises instructions which, when executed by one or more hardware processors, causes performance of any of the operations described herein and/or recited in any of the claims”.) Gama, Li, Xu, Dwork, and Deosthali are considered analogous to the claimed invention because they deal with data drift. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama,Li, Xu, and Dwork to be able to run on a processor and non-transitory computer-readable medium. One of the ordinary skill in the art would have known to apply the known technique of using processors and a non-transitory computer-readable medium to perform Gama and Li’s techniques. Therefore, applying Li’s technique would yield the predictable result of running instructions on processors and a non-transitory computer-readable medium (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results. Regarding claim 14, Gama, Li, Xu, Dwork, and Deosthali teach the non-transitory computer-readable medium of claim 13, Gama further teaches in response to determining that the first data drift exceeds the threshold, updating a machine learning model based on the second snapshot to generate a first updated machine learning model, wherein the machine learning model was previously trained on the first snapshot (Gama, page 44:18, section 3.2.3, “Large enough and distinct enough are defined by the Hoeffding bound, testing whether the average of the two (sub)windows is larger than ε-cut” and “Retraining approaches need some data buffer to be stored in memory. Retraining has been used to emulate incremental learning with batch-learning algorithms [Gama et al. 2004]. At the beginning, a model is trained with all the available data. Next, whenever new data arrives, the previous model is discarded, the new data is merged with the previous data, and a new model is learned on this data” (Gama, page 44:20, section 3.3.1). Examiner notes that the Hoeffding bound is being mapped to how the first data drift is determined). Regarding claim 15, Gama, Li, Xu, Dwork, and Deosthali teach the non-transitory computer-readable medium of claim 13, Gama further teaches in response to determining that the second data drift exceeds the threshold, updating the first updated machine learning model based on the third snapshot to generate a second updated machine learning model (Gama, page 44:20, section 3.3.1, “Retraining approaches need some data buffer to be stored in memory. Retraining has been used to emulate incremental learning with batch-learning algorithms [Gama et al. 2004]. At the beginning, a model is trained with all the available data. Next, whenever new data arrives, the previous model is discarded, the new data is merged with the previous data, and a new model is learned on this data”). Gama does not teach, but Deosthali does teach subsequent to the second updated machine learning model being generated: determining whether a performance metric of the second updated machine learning model exceeds a performance metric of the speculative machine learning model (Deosthali, page 6, paragraph 0058, “different models can be compared (e.g., A/B-tested) based on their respective accuracy and performance to identify the best model across different model attributes and data distributions”) in response to determining that the performance metric of the second updated machine learning model does not exceed the performance metric of the speculative machine learning model, relabeling the speculative machine learning model to be the second updated machine learning model (Desothali, page 6, paragraph 0059, “At block 360, the process 300 selects a machine learning for deployment. User can save the best model to create the baseline for further data and model experiments. Saving the base model saves the model attributes, hyperattributes, runtime information, and the model file for scoring new datasets. A most optimized model that accommodates changes in data with robust model metrics across wide ranges of data and with satisfactory non-functional (latency, cost) performance metrics can be chosen for deployment”. Examiner notes that “saving” the model is mapped to relabeling the model). in response to determining that the performance metric of the second updated machine learning model exceeds the performance metric of the speculative machine learning model, deploying the second updated machine learning model to replace the speculative machine learning model (Deosthali, page 2, paragraph 0027, “Based on the evaluation information, the machine learning system 105 can analyze the effect of drift and select best-performing machine learning model 115 for deployment to the production system 107 from the baseline model or the candidate models”). Gama, Li, Xu, Dwork, and Deosthali are considered analogous to the claimed invention because they deal with data drift. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama, Li, Xu, and Dwork to determine which machine learning model to use from Deosthali. One of the ordinary skill in the art would have known to apply the known technique of comparing machine learning models based on accuracy and performance. Therefore, applying Li’s technique would yield the predictable result of determining the best performing machine learning model to use (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results. Regarding claim 16, Gama, Li, Xu, Dwork, and Deosthali teach the non-transitory computer-readable medium of claim 13, Gama further teaches comparing the first data profile to the second data profile to process the first data drift using a detection algorithm, wherein the first data profile comprises a first plurality of metrics, wherein the first plurality of metrics comprises statistics of the first data profile, wherein the second data profile comprises a second plurality of metrics, and wherein the second plurality of metrics comprises statistics of the second data profile (Gama, page 44:18, section 3.2.3, “Whenever two large enough (sub)windows of W exhibit distinct enough means, the algorithm concludes that the expected values within those windows are different, and the older (sub)window is dropped. Large enough and distinct enough are defined by the Hoeffding bound, testing whether the average of the two (sub)windows is larger than ε-cut”. Examiner notes that the Hoeffding bound is the detecting algorithm). Gama does not teach, but Deosthali does teach comparing the first data profile to the second data profile to process the first data drift using a detection algorithm, wherein the first data profile comprises a first plurality of metrics, wherein the first plurality of metrics comprises statistics of the first data profile, wherein the second data profile comprises a second plurality of metrics, and wherein the second plurality of metrics comprises statistics of the second data profile (Deosthali, page 7, paragraph 0064, “the data validator 227 can generate a list of metrics of divergent dataset 215 attributes, such as distribution, mean, median mode, standard deviation, quantities of values in the dataset, and missing values, observations within one standard deviations (%), observations within two standard deviations (%), observations within three standard deviations (%), negative outliers (%), positive outliers (%), missing values (%), minimum value, maximum value, negative value, unique values, and zero values”). Gama, Li, Xu, Dwork, and Deosthali are considered analogous to the claimed invention because they deal with data drift. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama, Li, Xu, and Dwork to include a plurality of metrics from Deosthali to compare the sub windows from Gama. One of the ordinary skill in the art would have known to apply the known technique of using multiple metrics to compare datasets. Therefore, applying Li’s technique would yield the predictable result of accurately comparing datasets based on multiple variables (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results. Gama and Deosthali do not teach, but Li does teach determining the predicted data drift using a prediction machine learning model, wherein the prediction machine learning model is trained using a similarity between the first plurality of metrics and the second plurality of metrics, and wherein the prediction machine learning model is validated using the similarity between the second plurality of metrics and a third plurality of metrics, wherein the third plurality of metrics comprises statistics of the third data profile (Li, Abstract, “We first train a predictor to estimate the future data distribution, then leverage it to generate training samples, and finally train models on the generated data” and ” DDG-DA generates the resampling probability of each historical data sample to construct the future data distribution in estimation. However, it is quite challenging in reality to train this data generator to maximize the similarity between the predicted data distribution (represented by weighted resampling on historical data) and the ground truth future data distribution (represented by data in the future)” (Li, Abstract)). Gama, Xu, Dwork, and Deosthali are considered analogous to the claimed invention because they deal with data drift. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama, Xu, Dwork, and Deosthali to include Li’s prediction machine learning model. One of the ordinary skill in the art would have known to apply the known technique of predicting data based on past data. Therefore, applying Li’s technique would yield the predictable result of accurately predict future data (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results. Regarding claim 17, Gama, Li, Xu, Dwork, and Deosthali teach the non-transitory computer-readable medium of claim 13, Gama further teaches determining a performance threshold, wherein the performance threshold is related to a performance metric of a machine learning model (Gama, page 44:16, section 3.2.2, “A commonly used confidence level for Warning is 95% with the threshold pi + σi ≥ pmin + 2 ∗ σmin, and for Out-of-Control is 99% with the threshold pi +σi ≥ pmin +3∗σmin”. Examiner notes that pi +σi ≥ pmin +3∗σmin is mapped to the performance threshold where pi +σi is that performance metric and pmin +3∗σmin is that threshold”). in response to determining the machine learning model is not above the performance threshold, updating the machine learning model based on a new snapshot from the data stream (Gama, page 44:20, section 3.3.1, “Retraining has been used to emulate incremental learning with batch-learning algorithms [Gama et al. 2004]. At the beginning, a model is trained with all the available data. Next, whenever new data arrives, the previous model is discarded, the new data is merged with the previous data, and a new model is learned on this data”). Regarding claim 18, Gama, Li, Xu, Dwork, and Deosthali teach the non-transitory computer-readable medium of claim 13, Gama, Li, Dwork, and Deosthali do not teach, but Xu teaches generating, based on the predicted data drift, a synthetic snapshot of the data stream corresponding to the third time in the future further comprises: identifying a feature to replace in the second snapshot based on the predicted data drift (Xu, page 4, section 4.3, PNG media_image1.png 254 564 media_image1.png Greyscale Examiner notes that the discrete is being mapped to the feature). generating synthetic data to replace the feature in the second snapshot, wherein the synthetic data has a distribution of values that are similar to the second snapshot and are able to generate the predicted data drift (Xu, page 4, section 4.3, “Integrating a conditional generator into the architecture of a GAN requires to deal with the following issues: 1) it is necessary to devise a representation for the condition as well as to prepare an input for it, 2) it is necessary for the generated rows to preserve the condition as it is given, and 3) it is necessary for the conditional generator to learn the real data conditional distribution, i.e. PG(row|Di∗ = k∗) = P(row|Di∗ = k∗), so that we can reconstruct the original distribution as P(row) = Σ PG(row|Di∗ = k∗)P(Di∗ = k)”. Examiner notes that step 2 shows that the rows generated must be preserved or similar to the original data). replacing the feature in the second snapshot with the synthetic data (Xu, page 2, section 3, “After training [data synthesizer] G on [training set on table T] Ttrain, [synthetic table] Tsyn is constructed by independently sampling rows using G”). Gama, Li, Xu, Dwork, and Deosthali are considered analogous to the claimed invention because they deal with synthetic data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama, Li, Dwork, and Deosthali to include Xu’s conditional tabular generative adversarial network (CTGAN) method. Xu teaches that “CTGAN outperforms Bayesian methods on most of the real datasets whereas other deep learning methods could not” (Xu, page 1, Abstract) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Regarding claim 19, Gama, Li, Xu, Dwork, and Deosthali teach the non-transitory computer-readable medium of claim 13, Gama, Xu, Dwork, and Deosthali do not teach, but Li teaches iteratively sampling from the first snapshot and the second snapshot to generate synthetic data for the predicted data drift (Li, section 1, “DDG-DA generates the resampling probability of each historical data sample to construct the future data distribution in estimation … the data comes in a sequential mode, meaning that, at any timestamp t, the learning task can only observe the new coming information, i.e., Data(t), in addition to historical ones, i.e., Data(t−1), Data(t−2), ···, Data(1)”). combining the synthetic data to generate a synthetic snapshot (Li, section 1, “DDG-DA generates the resampling probability of each historical data sample to construct the future data distribution in estimation”). Gama, Li, Xu, Dwork, and Deosthali are considered analogous to the claimed invention because they deal with synthetic data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gama, Xu, Dwork, and Deosthali to include Li’s DDG-DA method to iterate through a dataset. Li teaches that “Only a few adaptation methods are model agnostic. DDG-DA focuses on predicting future data distribution, which is a model-agnostic solution and can benefit different customized forecasting models on streaming data.” (Li, section 2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. Regarding claim 20, Gama, Li, Xu, Dwork, and Deosthali teach the non-transitory computer-readable medium of claim 13, Gama further teaches the threshold is determined based on size or frequency of a data profile, and wherein determining the first data drift exceeds the threshold is based on a similarity between the first data profile and the second data profile (Gama, page 44:18, section, “Whenever two large enough (sub)windows of W exhibit distinct enough means, the algorithm concludes that the expected values within those windows are different, and the older (sub)window is dropped. Large enough and distinct enough are defined by the Hoeffding bound, testing whether the average of the two (sub)windows is larger than εcut”. Examiner notes that large enough is mapped to the size of the data while distinct enough is mapped to the similarity between the data). Response to Arguments On page 1-2, Applicant argues: Claims 1-20 are rejected under 35 U.S.C. § 101 as allegedly being directed to non- statutory subject matter. For at least the reasons agreed during the interview and the following reasons, and without acquiescing in the rejection, the amended independent claims, and the claims that depend thereon, are patent-eligible under 35 U.S.C. § 101. Even assuming the amended independent claims can reasonably be determined to recite an abstract idea under Prong One of Step 2A, which Applicant does not concede, MPEP 2106.04(d) states: "[a] claim reciting a judicial exception is not directed to the judicial exception if it also recites additional elements demonstrating that the claim as a whole integrates the exception into a practical application. One way to demonstrate such integration is when the claimed invention improves the functioning of a computer or improves another technology or technical field." The USPTO's December 4, 2025 Memorandum regarding the In re DeJardins decision states: "rejection was vacated because the claims were directed to training a machine learning model on multiple tasks, while preserving prior tasks performed, which properly integrated an otherwise abstract idea into a practical application; therefore, the claim satisfied Step 2 of the framework set forth in Alice Corp. v. CLS Bank Int'l, 573 U.S. 208 (2014). We specifically credited the claims for improving the functioning of the machine learning model itself." Representative amended claim 1 as a whole integrates into a practical application because it is directed to improvements in the technical field of machine learning. The claimed invention improves the performance of machine learning models by generating a speculative machine learning model for any predicted data drifts based on the previous data drifts that occurred in a manner that increases performance and reduces use of processing resources relative to conventional approaches. See, for example, paragraphs 1-5 of the specification. Specifically, the claim features of (i) extrapolating the first data drift to determine a predicted data drift corresponding to a third time in the future and generating, based on the predicted data drift, a synthetic snapshot of the data stream corresponding to the third time in the future; (ii) generating a speculative machine learning model from updating the first updated machine learning model based on the synthetic snapshot; (iii) determining, based on the second data profile and the third data profile, that a second data drift between the second snapshot and the third snapshot exceeds the threshold; and (iv) in response to determining that the second data drift exceeds the threshold, deploying the speculative machine learning model-solve the technical problem of how to project data drifts. Solving this technical problem provides the practical benefit of generating speculative machine learning models for future data drifts. Therefore, the above features of representative claim 1 provide improvements to the technical field of machine learning. Thus, the claims integrate the alleged abstract idea into a practical application and are patent-eligible under 35 U.S.C. § 101 for at least some of the reasons described above with regards to claim 1. Accordingly, Applicant respectfully requests that the Examiner reconsider and withdraw this rejection. Applicant’s argument that the claims provide an improvement to a technology or technical field, Examiner agrees. Paragraphs 5 and 6 of the Rejection under 35 U.S.C. § 101 section supports the improvement of projecting data drifts. Therefore, the 35 U.S.C. § 101 rejection has been withdrawn. On page 3-4, Applicant argues: Claims 3, 5, and 7-9 are rejected under 35 U.S.C. § 103 as allegedly being unpatentable over GAMA (A Survey on Concept Drift Adaptation) and LI (DDG-DA: Data Distribution Generation for Predictable Concept Drift Adaptation), and claims 1, 2, 4, 6, 13-17, and 20 are rejected under 35 U.S.C. § 103 as allegedly being unpatentable over GAMA, LI, and DEOSTHALI (US Patent Application Publication No. 2024/0086762, and claims 10 and 11 are rejected under 35 U.S.C. § 103 as allegedly being unpatentable over GAMA, LI, and XU (Modeling Tabular Data Using Conditional GAN), and claim 12 is rejected under 35 U.S.C. § 103 as allegedly being unpatentable over GAMA, LI, and DWORK (Calibrating Noise to Sensitivity in Private Data Analysis), and claims 18 and 19 are rejected under 35 U.S.C. § 103 as allegedly being unpatentable over GAMA, LI, DEOSTHALI, and XU. For at least the reasons agreed to during the interview and without acquiescing in the rejection, the cited sections of the applied references, whether taken alone or in any reasonable combination, do not disclose one or more features recited in the amended independent claims. Therefore, the amended independent claims and the claims that depend thereon are patentable over the cited sections of the applied references, whether taken alone or in any reasonable combination. Accordingly, Applicant respectfully requests that the Examiner reconsider and withdraw this rejection. Regarding the Applicant’s argument that the cited combination of references fail to disclose or suggest the amended claims, the Examiner respectfully disagrees. In claim 1, Li teaches iteratively sampling from the first snapshot of the data stream and the second snapshot of the data stream (Li, section 1, “DDG-DA generates the resampling probability of each historical data sample to construct the future data distribution in estimation … the data comes in a sequential mode, meaning that, at any timestamp t, the learning task can only observe the new coming information, i.e., Data(t), in addition to historical ones, i.e., Data(t−1), Data(t−2), ···, Data(1)”). Li’s DDG-DA has it’s Data(t) split into different timestamps t In claim 1, Xu teaches by one or more of replacing a feature, in the second snapshot of the data stream, with synthetic data (Xu, page 2, section 3, “After training [data synthesizer] G on [training set on table T] Ttrain, [synthetic table] Tsyn is constructed by independently sampling rows using G” and (Xu, page 4, section 4.3) PNG media_image1.png 254 564 media_image1.png Greyscale Examiner notes that the discrete column is being mapped to the feature). Xu shows that the CTGAN model generates synthetic data Tsyn based on the training set Ttrain on discrete column(s). In claim 1, Dworks teaches adding random noise, to the synthetic snapshot, based on an amount of random noise required based on the first data profile and the second data profile (Dwork, page 265, section 1.1, “we analyze the sensitivity of specific data analysis functions, including histograms, contingency tables, and covariance matrices, all of which have very high-dimensional output, and show that their sensitivities are independent of the dimension. Previous privacy-preserving approximations to these quantities used noise proportional to the dimension; the new analysis permits noise of size O(1)” and “the true answer is perturbed by the addition of random noise generated according to a carefully chosen distribution, and this response, the true answer plus noise, is returned to the user” (Dwork, page 265, section 1)). Dworks shows the addition of random noise to a dataset. Claims 3 and 13 follow the same reason as claim 1 since they are the method claim and non-transitory computer readable medium claim, respectfully. Examiner respectfully directs the Applicant to the above 103 rejection. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Gama et al. (“Learning with Drift Detection”) discloses a method for detection of changes in the probability distribution of examples. Chawla et al. ("SMOTE: Synthetic Minority Over-sampling Technique") discloses techniques for generating synthetic minority class samples by iteratively selecting each minority class data points. Wenchel et al. (US 11922280 B2) discloses monitoring performance on a machine learning system based on a plurality of metrics. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to STEVEN VO whose telephone number is (571)272-9622. The examiner can normally be reached Monday - Friday from 7-3 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /S.V./Examiner, Art Unit 2148 /MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148
Read full office action

Prosecution Timeline

Aug 22, 2023
Application Filed
May 05, 2026
Non-Final Rejection mailed — §101, §103
Jul 14, 2026
Applicant Interview (Telephonic)
Jul 15, 2026
Response Filed
Jul 20, 2026
Examiner Interview Summary
Sep 24, 2026
Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
Grant Probability
Moderate
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month