DETAILED ACTION
This action is in response to the application filed on 08/14/2023. Claims 1-20 are pending and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 1, 12-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gu et al. (Empirical Asset Pricing via Machine Learning) (hereafter referred to as Gu) in view of Schwiep et al. (US 20220292308 A1) (hereafter referred to as Schwiep) and Lopez de Prado et al. (Advances in Financial Machine Learning) (hereafter referred to as Lopez) .
Regarding claim 1:
Gu teaches
determining, from received panel data, a variable to uniquely identify records along a cross-sectional dimension in the panel data (Gu, page 4, section 1.5, “We conduct a large scale empirical analysis, investigating nearly 30,000 individual stocks over 60 years from 1957 to 2016. Our predictor set includes 94 characteristics for each stock, interactions of each characteristic with eight aggregate time series variables, and 74 industry sector dummy variables, totaling more than 900 baseline signals” and “In its most general form, we describe an asset’s excess return as an additive prediction error model: ri,t+1 = Et(ri,t+1) + i,t+1, where Et(ri,t+1) = g (zi,t). (1) (2) Stocks are indexed as i = 1,...,Nt and months by t = 1,...,T. For ease of presentation, we assume a balanced panel of stocks, and defer the discussion on missing data to Section 3.1” (Gu, page 8, section 2). Examiner notes that the multiple stocks being observed over a period of time is mapped to panel data and the stocks being indexed as i shows indicates that i is the unique variable).
splitting, based on the variable, the panel data into a train dataset, a test dataset, and a validation dataset (Gu, page 4, section 1.5, “We conduct a large scale empirical analysis, investigating nearly 30,000 individual stocks over 60 years from 1957 to 2016. Our predictor set includes 94 characteristics for each stock, interactions of each characteristic with eight aggregate time series variables, and 74 industry sector dummy variables, totaling more than 900 baseline signals” and “we divide our sample into three disjoint time periods that maintain the temporal ordering of the data. The first, or “training,” subsample is used to estimate the model subject to a specific set of tuning parameter values. The second, or “validation,” sample is used for tuning the hyperparameters … the third, or “testing,” subsample, which is used for neither estimation nor tuning, is truly out-of-sample and thus is used to evaluate a method’s predictive performance” (Gu, page 9, section 2.1). Examiner notes that the multiple stocks being observed over a period of time is mapped to panel data).
determining, based on the validation dataset, one or more time intervals and removing, from the validation dataset, data falling outside of the one or more time intervals, wherein the validation dataset includes data that is out-of-sample and out-of-time with respect to the train dataset (Gu , page 4, section 1.5, “The second, or “validation,” sample is used for tuning the hyperparameters”).
training a machine learning model based on the train dataset, validating the machine learning model on the validation dataset, and testing the machine learning model on the test dataset, wherein the validation dataset is used to generate an estimate of model error that is out-of-sample and out-of-time with respect to the train dataset (Gu, page 9, section 2.1, “We follow the most common approach in the literature and select tuning parameters adaptively from the data in a validation sample. In particular, we divide our sample into three disjoint time periods that maintain the temporal ordering of the data. The first, or “training,” subsample is used to estimate the model subject to a specific set of tuning parameter values. The second, or “validation,” sample is used for tuning the hyperparameters. We construct forecasts for data points in the validation sample based on the estimated model from the training sample. Next, we calculate the objective function based on forecast errors from the validation sample, and iteratively search for hyperparameters that optimize the validation objective (at each step re-estimating the model from the training data subject to the prevailing hyperparameter values). Thus the third, or “testing,” subsample, which is used for neither estimation nor tuning, is truly out-of-sample and thus is used to evaluate a method’s predictive performance”. Also from Gu, page 25, section 3.1, “We divide the 60 years of data into 18 years of training sample (1957- 1974), 12 years of validation sample (1975- 1986), and the remaining 30 years (1987- 2016) for out-of-sample testing” and “Table 1 presents the comparison of machine learning techniques in terms of their out-of-sample predictive R2”( Gu, page 25, section 3.2,). Examiner notes that the training set is out-of-time compared to the validation set).
generating, using the machine learning model, an output based on new input (Gu, page 6, section 1.6, “Machine learning has great potential for improving risk premium measurement, which is fundamentally a problem of prediction. It amounts to best approximating the conditional expectation E(ri,t+1|Ft), where ri,t+1 is an asset’s return in excess of the risk-free rate, and Ft is the true and unobservable information set of market participants”. Examiner notes that the conditional expectation is mapped to the output, where t+1 is that new input).
Gu does not teach, but Lopez does teach
removing, from the validation dataset and after splitting the panel data into the train dataset, the test dataset, and the validation dataset, data falling within an out- of-time period that is after a particular period of time of the panel data (Lopez, Chapter 7.4.2, “The embargo does not need to affect training observations prior to a test set, because training labels Yi = f[[ti, 0, ti, 1]], where ti, 1 < tj, 0 (training ends before testing begins), contain information that was available at the testing time tj, 0. In other words, we are only concerned with training labels Yi = f[[ti, 0, ti, 1]] that take place immediately after the test, tj, 1 ≤ ti, 0 ≤ tj, 1 + h.” and “In the previous sections we have discussed how to produce training/testing splits when labels overlap. That introduced the notion of purging and embargoing, in the particular context of model development. In general, we need to purge and embargo overlapping training observations whenever we produce a train/test split, whether it is for hyper-parameter fitting, backtesting, or performance evaluation” (Lopez, Chapter 7.4.3)
PNG
media_image1.png
311
633
media_image1.png
Greyscale
Examiner notes that we can apply this embargo function to the validation dataset mentioned above).
removing, from the train dataset, data falling within the out-of-time period (Lopez, Chapter 7.4.2, “The embargo does not need to affect training observations prior to a test set, because training labels Yi = f[[ti, 0, ti, 1]], where ti, 1 < tj, 0 (training ends before testing begins), contain information that was available at the testing time tj, 0. In other words, we are only concerned with training labels Yi = f[[ti, 0, ti, 1]] that take place immediately after the test, tj, 1 ≤ ti, 0 ≤ tj, 1 + h.”)
PNG
media_image1.png
311
633
media_image1.png
Greyscale
determining, based on the validation dataset, one or more time intervals and removing, from the validation dataset, data falling outside of the one or more time intervals, wherein the validation dataset includes data that is out-of-sample and out-of-time with respect to the train dataset (Lopez, chapter 7.4, “One way to reduce leakage is to purge from the training set all observations whose labels overlapped in time with those labels included in the testing set. I call this process “purging”.
PNG
media_image2.png
304
630
media_image2.png
Greyscale
Examiner notes that this function drops (removes) training data that fall within a time interval and we can apply this purging function to the validation set mentioned above).
removing, from the train dataset, data falling within the one or more time intervals, wherein the train dataset includes data that is out-of-sample and out-of-time with respect to the validation dataset (Lopez, chapter 7.4, “One way to reduce leakage is to purge from the training set all observations whose labels overlapped in time with those labels included in the testing set. I call this process “purging”.
PNG
media_image2.png
304
630
media_image2.png
Greyscale
Examiner notes that this function drops (removes) training data that fall within a time interval).
Gu and Lopez are considered analogous to the claimed invention because they both deal with time-series data and cross-sectional data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gu to use Lopez’s purging and embargo function to filter data out from Gu’s training and validation datasets depending on the out-of-time period or time interval(s). One of the ordinary skill in the art would have known to apply the known technique of using the purging and embargo functions to remove data based on one or more time periods. Therefore, applying Lopez’s technique would yield the predictable result of keeping datasets out-of-time from each other to prevent data leakage (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results.
Gu and Lopez does not teach, but Schweip does teach
one or more processors (Schweip, paragraph 0156, “The term “system” may encompass all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers”).
one or more non-transitory computer-readable media storing instructions that when executed by the one or more processors cause operations (Schweip, paragraph 0150, “In some implementations, at least a portion of the approaches described above may be realized by instructions that upon execution cause one or more processing devices to carry out the processes and functions described above. Such instructions may include, for example, interpreted instructions such as script instructions, or executable code, or other instructions stored in a non-transitory computer readable medium”).
Gu, Lopez, and Schweip are considered analogous to the claimed invention because they all deal with time-series data. It would have been obvious to one having ordinary skill in the art prior to the effective filling date to have modified Gu and Lopez to include one or more processors as well as one or more non-transitory computer-readable mediums. One of the ordinary skill in the art would have known to apply Schweip’s technique of using a processor and a non-transitory computer readable medium to perform Gu and Lopez’s technique. Therefore, applying Schweip’s technique would have yielded predictable results of running instructions on a computer (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predictable results.
Regarding claim 12:
Gu teaches
splitting panel data into a train dataset, a test dataset, and a validation dataset (Gu, page 4, section 1.5, “We conduct a large scale empirical analysis, investigating nearly 30,000 individual stocks over 60 years from 1957 to 2016. Our predictor set includes 94 characteristics for each stock, interactions of each characteristic with eight aggregate time series variables, and 74 industry sector dummy variables, totaling more than 900 baseline signals” and “we divide our sample into three disjoint time periods that maintain the temporal ordering of the data. The first, or “training,” subsample is used to estimate the model subject to a specific set of tuning parameter values. The second, or “validation,” sample is used for tuning the hyperparameters … the third, or “testing,” subsample, which is used for neither estimation nor tuning, is truly out-of-sample and thus is used to evaluate a method’s predictive performance” (Gu, page 9, section 2.1). Examiner notes that the multiple stocks being observed over a period of time is mapped to panel data).
Gu does not teach, but Lopez does teach
removing, from the validation dataset and after splitting the panel data into the train dataset, the test dataset, and the validation dataset, data falling within an out- of-time period that is after a particular period of time of the panel data (Lopez, Chapter 7.4.2, “The embargo does not need to affect training observations prior to a test set, because training labels Yi = f[[ti, 0, ti, 1]], where ti, 1 < tj, 0 (training ends before testing begins), contain information that was available at the testing time tj, 0. In other words, we are only concerned with training labels Yi = f[[ti, 0, ti, 1]] that take place immediately after the test, tj, 1 ≤ ti, 0 ≤ tj, 1 + h.” and “In the previous sections we have discussed how to produce training/testing splits when labels overlap. That introduced the notion of purging and embargoing, in the particular context of model development. In general, we need to purge and embargo overlapping training observations whenever we produce a train/test split, whether it is for hyper-parameter fitting, backtesting, or performance evaluation” (Lopez, Chapter 7.4.3)
PNG
media_image1.png
311
633
media_image1.png
Greyscale
Examiner notes that we can apply this embargo function to the validation dataset mentioned above).
removing, from the train dataset, data falling within the out-of-time period (Lopez, Chapter 7.4.2, “The embargo does not need to affect training observations prior to a test set, because training labels Yi = f[[ti, 0, ti, 1]], where ti, 1 < tj, 0 (training ends before testing begins), contain information that was available at the testing time tj, 0. In other words, we are only concerned with training labels Yi = f[[ti, 0, ti, 1]] that take place immediately after the test, tj, 1 ≤ ti, 0 ≤ tj, 1 + h.”)
PNG
media_image1.png
311
633
media_image1.png
Greyscale
determining one or more time intervals and removing, from the validation dataset, data falling outside of the one or more time intervals, wherein the validation dataset includes data that is out-of-sample and out-of-time with respect to the train dataset (Lopez, chapter 7.4, “One way to reduce leakage is to purge from the training set all observations whose labels overlapped in time with those labels included in the testing set. I call this process “purging”.
PNG
media_image2.png
304
630
media_image2.png
Greyscale
Examiner notes that this function drops (removes) training data that fall within a time interval and we can apply this purging function to the validation set mentioned above).
removing, from the train dataset, data falling within the one or more time intervals (Lopez, chapter 7.4, “One way to reduce leakage is to purge from the training set all observations whose labels overlapped in time with those labels included in the testing set. I call this process “purging”.
PNG
media_image2.png
304
630
media_image2.png
Greyscale
Examiner notes that this function drops (removes) training data that fall within a time interval) removing, from the train dataset, data falling within the one or more time intervals.
Gu and Lopez are considered analogous to the claimed invention because they both deal with time-series data and cross-sectional data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gu to use Lopez’s purging and embargo function to filter data out from Gu’s training and validation datasets depending on the out-of-time period or time interval(s). One of the ordinary skill in the art would have known to apply the known technique of using the purging and embargo functions to remove data based on one or more time periods. Therefore, applying Lopez’s technique would yield the predictable result of keeping datasets out-of-time from each other to prevent data leakage (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results.
Gu and Lopez does not teach, but Schweip does teach
One or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors cause operations (Schweip, paragraph 0150, “In some implementations, at least a portion of the approaches described above may be realized by instructions that upon execution cause one or more processing devices to carry out the processes and functions described above. Such instructions may include, for example, interpreted instructions such as script instructions, or executable code, or other instructions stored in a non-transitory computer readable medium”).
Gu, Lopez, and Schweip are considered analogous to the claimed invention because they all deal with time-series data. It would have been obvious to one having ordinary skill in the art prior to the effective filling date to have modified Gu and Lopez to include one or more processors as well as one or more non-transitory computer-readable mediums. One of the ordinary skill in the art would have known to apply Schweip’s technique of using a processor and a non-transitory computer readable medium to perform Gu and Lopez’s technique. Therefore, applying Schweip’s technique would have yielded predictable results of running instructions on a computer (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predictable results.
Regarding claim 13:
Gu, Lopez, and Schweip teach the product of claim 12, Gu further teaches
determining a performance metric for a first machine learning model for the out-of-time period based on the test dataset (Gu, page 1, section 1.1, “Our primary contributions are two-fold. First, we provide a set of benchmarks for the predictive accuracy of machine learning methods ... The first is a comparatively high out-of-sample predictive R2 relative to preceding literature”. Examiner notes that the R2 model is being mapped to the performance metric).
in response to determining whether the performance metric is below a threshold, determining updated one or more time intervals (Gu, page 21, section 2.7, “Next, “early stopping” is a general machine learning regularization tool. It begins from an initial parameter guess that imposes parsimonious parameterization (for example, setting all θ values close to zero). In each step of the optimization algorithm, the parameter guesses are gradually updated to reduce prediction errors in the training sample. At each new guess, predictions are also constructed for the validation sample, and the optimization is terminated when the validation sample errors begin to increase”. Examiner notes that the parameter guesses are being mapped to the time intervals).
Gu does not teach, but Lopez does teach
removing from the train dataset and the validation dataset falling within the updated one or more time intervals (Lopez, chapter 7.4, “One way to reduce leakage is to purge from the training set all observations whose labels overlapped in time with those labels included in the testing set. I call this process “purging”.
PNG
media_image2.png
304
630
media_image2.png
Greyscale
Examiner notes that this function drops (removes) training data that fall within a time interval) removing, from the train dataset, data falling within the one or more time intervals). This function can be applied the validation set mentioned above.
Gu, Lopez, and Schweip are considered analogous to the claimed invention because they all deal with time-series data. It would have been obvious to one having ordinary skill in the art prior to the effective filling date to have modified Gu to include the purging function from Lopez when early stopping occurs. One of the ordinary skill in the art would have known to apply Lopez’s technique of removing data when new parameters are updated. Therefore, applying Lopez’s technique would have yielded the predictable result of improving performance by updating parameters and updating data to maintain accuracy (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predictable results.
Regarding claim 14:
Gu, Lopez, and Schweip teach the product of claim 13, Gu further teaches
generating the threshold based on the performance metric of the first machine learning model on the validation dataset (Gu, page 21, section 2.7, “Next, “early stopping” is a general machine learning regularization tool. It begins from an initial parameter guess that imposes parsimonious parameterization (for example, setting all θ values close to zero). In each step of the optimization algorithm, the parameter guesses are gradually updated to reduce prediction errors in the training sample. At each new guess, predictions are also constructed for the validation sample, and the optimization is terminated when the validation sample errors begin to increase”. Examiner notes that the parameter guesses are being mapped to the time intervals).
Regarding claim 15:
Gu, Lopez, and Schweip teach the product of claim 12, Gu further teaches
determining a performance metric for a first machine learning model (Gu, page 1, section 1.1, “Our primary contributions are two-fold. First, we provide a set of benchmarks for the predictive accuracy of machine learning methods ... The first is a comparatively high out-of-sample predictive R2 relative to preceding literature”. Examiner notes that the R2 model is being mapped to the performance metric).
in response to determining whether the performance metric is below a threshold, modifying a set of hyperparameters of the first machine learning model (Gu, page 21, section 2.7, “Next, “early stopping” is a general machine learning regularization tool. It begins from an initial parameter guess that imposes parsimonious parameterization (for example, setting all θ values close to zero). In each step of the optimization algorithm, the parameter guesses are gradually updated to reduce prediction errors in the training sample. At each new guess, predictions are also constructed for the validation sample, and the optimization is terminated when the validation sample errors begin to increase”).
Regarding claim 16:
Gu, Lopez, and Schweip teach the product of claim 12, Gu further teaches
receiving new panel data and determining a new variable to uniquely identify records along a cross-sectional dimension in the new panel data (Gu, page 64, Section D Sample Splitting, “A common alternative to the fixed split scheme is a “rolling” scheme, in which the training and validation samples gradually shift forward in time to include more recent data, but holds the total number of time periods in each training and validation sample fixed. For each rolling window, one re f its the model from the prevailing training and validation samples, and tracks a model’s performance in the remaining test data that has not been subsumed by the rolling windows” and “In its most general form, we describe an asset’s excess return as an additive prediction error model: ri,t+1 = Et(ri,t+1) + i,t+1, where Et(ri,t+1) = g (zi,t). (1) (2) Stocks are indexed as i = 1,...,Nt and months by t = 1,...,T. For ease of presentation, we assume a balanced panel of stocks, and defer the discussion on missing data to Section 3.1” (Gu, page 8, section 2). Examiner notes that the shifting in time means new data is being received).
splitting, based on the new variable, the new panel data into a new train dataset, a new test dataset, and a new validation dataset (Gu, page 4, section 1.5, “We conduct a large scale empirical analysis, investigating nearly 30,000 individual stocks over 60 years from 1957 to 2016. Our predictor set includes 94 characteristics for each stock, interactions of each characteristic with eight aggregate time series variables, and 74 industry sector dummy variables, totaling more than 900 baseline signals” and “we divide our sample into three disjoint time periods that maintain the temporal ordering of the data. The first, or “training,” subsample is used to estimate the model subject to a specific set of tuning parameter values. The second, or “validation,” sample is used for tuning the hyperparameters … the third, or “testing,” subsample, which is used for neither estimation nor tuning, is truly out-of-sample and thus is used to evaluate a method’s predictive performance” (Gu, page 9, section 2.1). Examiner notes that the multiple stocks being observed over a period of time is mapped to panel data).
determining, based on the new validation dataset, one or more time intervals and removing, from the new validation dataset, data falling outside of the one or more time intervals, wherein the new validation dataset includes data that is out-of-sample and out-of-time with respect to the new train dataset (Gu , page 4, section 1.5, “The second, or “validation,” sample is used for tuning the hyperparameters”).
training a second machine learning model based on the new train dataset (Gu, page 9, section 2.1, “We follow the most common approach in the literature and select tuning parameters adaptively from the data in a validation sample. In particular, we divide our sample into three disjoint time periods that maintain the temporal ordering of the data. The first, or “training,” subsample is used to estimate the model subject to a specific set of tuning parameter values).
Gu does not teach, but Lopez does teach
removing, from the validation dataset and after splitting the panel data into the train dataset, the test dataset, and the validation dataset, data falling within an out- of-time period that is after a particular period of time of the panel data (Lopez, Chapter 7.4.2, “The embargo does not need to affect training observations prior to a test set, because training labels Yi = f[[ti, 0, ti, 1]], where ti, 1 < tj, 0 (training ends before testing begins), contain information that was available at the testing time tj, 0. In other words, we are only concerned with training labels Yi = f[[ti, 0, ti, 1]] that take place immediately after the test, tj, 1 ≤ ti, 0 ≤ tj, 1 + h.” and “In the previous sections we have discussed how to produce training/testing splits when labels overlap. That introduced the notion of purging and embargoing, in the particular context of model development. In general, we need to purge and embargo overlapping training observations whenever we produce a train/test split, whether it is for hyper-parameter fitting, backtesting, or performance evaluation” (Lopez, Chapter 7.4.3)
PNG
media_image1.png
311
633
media_image1.png
Greyscale
Examiner notes that we can apply this embargo function to the validation dataset mentioned above).
removing, from the train dataset, data falling within the out-of-time period (Lopez, Chapter 7.4.2, “The embargo does not need to affect training observations prior to a test set, because training labels Yi = f[[ti, 0, ti, 1]], where ti, 1 < tj, 0 (training ends before testing begins), contain information that was available at the testing time tj, 0. In other words, we are only concerned with training labels Yi = f[[ti, 0, ti, 1]] that take place immediately after the test, tj, 1 ≤ ti, 0 ≤ tj, 1 + h.”)
PNG
media_image1.png
311
633
media_image1.png
Greyscale
determining, based on the new validation dataset, one or more time intervals and removing, from the new validation dataset, data falling outside of the one or more time intervals, wherein the new validation dataset includes data that is out-of-sample and out-of-time with respect to the new train dataset (Lopez, chapter 7.4, “One way to reduce leakage is to purge from the training set all observations whose labels overlapped in time with those labels included in the testing set. I call this process “purging”.
PNG
media_image2.png
304
630
media_image2.png
Greyscale
Examiner notes that this function drops (removes) training data that fall within a time interval and we can apply this purging function to the validation set mentioned above).
removing, from the new train dataset, data falling within the one or more time intervals, wherein the new train dataset includes data that is out-of-sample and out-of-time with respect to the new validation dataset (Lopez, chapter 7.4, “One way to reduce leakage is to purge from the training set all observations whose labels overlapped in time with those labels included in the testing set. I call this process “purging”.
PNG
media_image2.png
304
630
media_image2.png
Greyscale
Examiner notes that this function drops (removes) training data that fall within a time interval).
Gu and Lopez are considered analogous to the claimed invention because they both deal with time-series data and cross-sectional data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gu to use Lopez’s purging and embargo function to filter data out from Gu’s training and validation datasets depending on the out-of-time period or time interval(s). One of the ordinary skill in the art would have known to apply the known technique of using the purging and embargo functions to remove data based on one or more time periods. Therefore, applying Lopez’s technique would yield the predictable result keeping datasets out-of-time from each other to prevent data leakage (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results.
Regarding claim 17:
Gu, Lopez, and Schweip teach the product of claim 12, Gu further teaches
splitting the panel data comprises preprocessing the panel data to remove any instances of incomplete data (Gu, page 53, Section A Monte Carlo Simulations, “We simulate the panel of characteristics for each 1 ≤ i ≤ N and each 1 ≤ j ≤ Pc from the following model: cij,t = 2 N +1CSrank(¯cij,t) − 1, ¯cij,t = ρj¯cij,t−1 + ij,t, (A.1) where ρj ∼ U[0.9,1], ij,t ∼ N(0,1 − ρ2 j) and CSrank is Cross-Section rank function, so that the characteristics feature some degree of persistence over time, yet is cross-sectionally normalized to be within [−1,1]. This matches our data cleaning procedure in the empirical study”. Examiner notes that Gu’s data cleaning procedure normalizes the data, which includes removing incomplete data).
Regarding claim 18:
Gu, Lopez, and Schweip teach the product of claim 12, Gu further teaches
determining, from panel data, a variable to uniquely identify records along a cross-sectional dimension in the panel data (Gu, page 8, section 2, “In its most general form, we describe an asset’s excess return as an additive prediction error model: ri,t+1 = Et(ri,t+1) + i,t+1, where Et(ri,t+1) = g (zi,t). (1) (2) Stocks are indexed as i = 1,...,Nt and months by t = 1,...,T. For ease of presentation, we assume a balanced panel of stocks, and defer the discussion on missing data to Section 3.1”. Examiner notes that the multiple stocks being observed over a period of time is mapped to panel data and the stocks being indexed as i shows indicates that i is the unique variable).
determining, based on the panel data, a size for the train dataset, test dataset, and validation dataset, wherein the size corresponds to a time period of the panel data (Gu, page 25, section 3.1, “We divide the 60 years of data into 18 years of training sample (1957- 1974), 12 years of validation sample (1975- 1986), and the remaining 30 years (1987- 2016) for out-of-sample testing”)
Regarding claim 19:
Gu, Lopez, and Schweip teach the product of claim 12, Gu further teaches
the test dataset is used to generate an estimate of model error that is out-of-sample and out-of-time with respect to the validation dataset and the train dataset (Gu, page 25, section 3.1, “We divide the 60 years of data into 18 years of training sample (1957- 1974), 12 years of validation sample (1975- 1986), and the remaining 30 years (1987- 2016) for out-of-sample testing” and “Table 1 presents the comparison of machine learning techniques in terms of their out-of-sample predictive R2” (Gu, page 25, section 3.2). Examiner notes that the training sample is out-of-time compared to the validation sample and training sample).
Regarding claim 20:
Gu, Lopez, and Schweip teach the product of claim 16, Gu further teaches
comparing a first machine learning model to the second machine learning model based on performance metrics to determine which machine learning model performs better (Gu, page 26, table 1)
PNG
media_image3.png
543
813
media_image3.png
Greyscale
Examiner notes that the table shows the comparison based on the R2 model.
Gu does not teach, but Lopez does teach
in response to determining that the second machine learning model outperforms the first machine learning model, averaging outputs from the first machine learning model and the second machine learning model to generate a new output (Lopez, Chapter 6.3, “Bootstrap aggregation (bagging) is an effective way of reducing the variance in forecasts. It works as follows: First, generate N training datasets by random sampling with replacement. Second, fit N estimators, one on each training set. These estimators are fit independently from each other, hence the models can be fit in parallel. Third, the ensemble forecast is the simple average of the individual forecasts from the N models”).
Gu and Lopez are considered analogous to the claimed invention because they both deal with time-series data and cross-sectional data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gu to use Lopez’s ensemble forecast to find the average between the machine learning models from Gu. One of the ordinary skill in the art would have known to apply the known technique of averaging outputs of machine learning models. Therefore, applying Lopez’s technique would yield the predictable result of averaging two machine learning outputs (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results.
Claim(s) 2-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gu in view of Lopez.
Regarding claim 2:
Gu teaches
Splitting received panel data into a train dataset, a test dataset, and a validation dataset (Gu, page 4, section 1.5, “We conduct a large scale empirical analysis, investigating nearly 30,000 individual stocks over 60 years from 1957 to 2016. Our predictor set includes 94 characteristics for each stock, interactions of each characteristic with eight aggregate time series variables, and 74 industry sector dummy variables, totaling more than 900 baseline signals” and “we divide our sample into three disjoint time periods that maintain the temporal ordering of the data. The first, or “training,” subsample is used to estimate the model subject to a specific set of tuning parameter values. The second, or “validation,” sample is used for tuning the hyperparameters … the third, or “testing,” subsample, which is used for neither estimation nor tuning, is truly out-of-sample and thus is used to evaluate a method’s predictive performance” (Gu, page 9, section 2.1). Examiner notes that the multiple stocks being observed over a period of time is mapped to panel data).
determining, based on the validation dataset, one or more time intervals and removing, from the validation dataset, data falling outside of the one or more time intervals, wherein the validation dataset includes data that is out-of-sample and out-of-time with respect to the train dataset (Gu, page 4, section 1.5, “The second, or “validation,” sample is used for tuning the hyperparameters”).
training a machine learning model based on the train dataset, validating the machine learning model on the validation dataset, and testing the machine learning model on the test dataset, wherein the validation dataset is used to generate an estimate of model error that is out-of-sample and out-of-time with respect to the train dataset (Gu, page 9, section 2.1, “We follow the most common approach in the literature and select tuning parameters adaptively from the data in a validation sample. In particular, we divide our sample into three disjoint time periods that maintain the temporal ordering of the data. The first, or “training,” subsample is used to estimate the model subject to a specific set of tuning parameter values. The second, or “validation,” sample is used for tuning the hyperparameters. We construct forecasts for data points in the validation sample based on the estimated model from the training sample. Next, we calculate the objective function based on forecast errors from the validation sample, and iteratively search for hyperparameters that optimize the validation objective (at each step re-estimating the model from the training data subject to the prevailing hyperparameter values). Thus the third, or “testing,” subsample, which is used for neither estimation nor tuning, is truly out-of-sample and thus is used to evaluate a method’s predictive performance”. Also from Gu, page 25, section 3.1, “We divide the 60 years of data into 18 years of training sample (1957- 1974), 12 years of validation sample (1975- 1986), and the remaining 30 years (1987- 2016) for out-of-sample testing” and “Table 1 presents the comparison of machine learning techniques in terms of their out-of-sample predictive R2”( Gu, page 25, section 3.2,). Examiner notes that the training set is out-of-time compared to the validation set).
Gu does not teach, but Lopez does teach
removing, from the validation dataset and after splitting the panel data into the train dataset, the test dataset, and the validation dataset, data falling within an out- of-time period that is after a particular period of time of the panel data (Lopez, Chapter 7.4.2, “The embargo does not need to affect training observations prior to a test set, because training labels Yi = f[[ti, 0, ti, 1]], where ti, 1 < tj, 0 (training ends before testing begins), contain information that was available at the testing time tj, 0. In other words, we are only concerned with training labels Yi = f[[ti, 0, ti, 1]] that take place immediately after the test, tj, 1 ≤ ti, 0 ≤ tj, 1 + h.” and “In the previous sections we have discussed how to produce training/testing splits when labels overlap. That introduced the notion of purging and embargoing, in the particular context of model development. In general, we need to purge and embargo overlapping training observations whenever we produce a train/test split, whether it is for hyper-parameter fitting, backtesting, or performance evaluation” (Lopez, Chapter 7.4.3)
PNG
media_image1.png
311
633
media_image1.png
Greyscale
Examiner notes that we can apply this embargo function to the validation dataset mentioned above).
removing, from the train dataset, data falling within the out-of-time period (Lopez, Chapter 7.4.2, “The embargo does not need to affect training observations prior to a test set, because training labels Yi = f[[ti, 0, ti, 1]], where ti, 1 < tj, 0 (training ends before testing begins), contain information that was available at the testing time tj, 0. In other words, we are only concerned with training labels Yi = f[[ti, 0, ti, 1]] that take place immediately after the test, tj, 1 ≤ ti, 0 ≤ tj, 1 + h.”)
PNG
media_image1.png
311
633
media_image1.png
Greyscale
determining, based on the validation dataset, one or more time intervals and removing, from the validation dataset, data falling outside of the one or more time intervals, wherein the validation dataset includes data that is out-of-sample and out-of-time with respect to the train dataset (Lopez, chapter 7.4, “One way to reduce leakage is to purge from the training set all observations whose labels overlapped in time with those labels included in the testing set. I call this process “purging”.
PNG
media_image2.png
304
630
media_image2.png
Greyscale
Examiner notes that this function drops (removes) training data that fall within a time interval and we can apply this purging function to the validation set mentioned above).
removing, from the train dataset, data falling within the one or more time intervals, wherein the train dataset includes data that is out-of-sample and out-of-time with respect to the validation dataset (Lopez, chapter 7.4, “One way to reduce leakage is to purge from the training set all observations whose labels overlapped in time with those labels included in the testing set. I call this process “purging”.
PNG
media_image2.png
304
630
media_image2.png
Greyscale
Examiner notes that this function drops (removes) training data that fall within a time interval).
Gu and Lopez are considered analogous to the claimed invention because they both deal with time-series data and cross-sectional data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gu to use Lopez’s purging and embargo function to filter data out from Gu’s training and validation datasets depending on the out-of-time period or time interval(s). One of the ordinary skill in the art would have known to apply the known technique of using the purging and embargo functions to remove data based on one or more time periods. Therefore, applying Lopez’s technique would yield the predictable result of keeping datasets out-of-time from each other to prevent data leakage (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results.
Regarding claim 3:
Gu and Lopez teach the system of claim 2, Gu further teaches
determining a performance metric for a first machine learning model for the out-of-time period based on the test dataset (Gu, page 1, section 1.1, “Our primary contributions are two-fold. First, we provide a set of benchmarks for the predictive accuracy of machine learning methods ... The first is a comparatively high out-of-sample predictive R2 relative to preceding literature”. Examiner notes that the R2 model is being mapped to the performance metric).
in response to determining whether the performance metric is below a threshold, determining updated one or more time intervals (Gu, page 21, section 2.7, “Next, “early stopping” is a general machine learning regularization tool. It begins from an initial parameter guess that imposes parsimonious parameterization (for example, setting all θ values close to zero). In each step of the optimization algorithm, the parameter guesses are gradually updated to reduce prediction errors in the training sample. At each new guess, predictions are also constructed for the validation sample, and the optimization is terminated when the validation sample errors begin to increase”. Examiner notes that the parameter guesses are being mapped to the time intervals).
Gu does not teach, but Lopez does teach
removing from the train dataset and the validation dataset falling within the updated one or more time intervals (Lopez, chapter 7.4, “One way to reduce leakage is to purge from the training set all observations whose labels overlapped in time with those labels included in the testing set. I call this process “purging”.
PNG
media_image2.png
304
630
media_image2.png
Greyscale
Examiner notes that this function drops (removes) training data that fall within a time interval) removing, from the train dataset, data falling within the one or more time intervals). This function can be applied the validation set mentioned above.
Gu and Lopez are considered analogous to the claimed invention because they all deal with time-series data. It would have been obvious to one having ordinary skill in the art prior to the effective filling date to have modified Gu to include the purging function from Lopez when early stopping occurs. One of the ordinary skill in the art would have known to apply Lopez’s technique of removing data when new parameters are updated. Therefore, applying Lopez’s technique would have yielded the predictable result of improving performance by updating parameters and updating data to maintain accuracy (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predictable results.
Regarding claim 4:
Gu and Lopez teach the method of claim 3, Gu further teaches
generating the threshold based on the performance metric of the first machine learning model on the validation dataset (Gu, page 21, section 2.7, “Next, “early stopping” is a general machine learning regularization tool. It begins from an initial parameter guess that imposes parsimonious parameterization (for example, setting all θ values close to zero). In each step of the optimization algorithm, the parameter guesses are gradually updated to reduce prediction errors in the training sample. At each new guess, predictions are also constructed for the validation sample, and the optimization is terminated when the validation sample errors begin to increase”. Examiner notes that the parameter guesses are being mapped to the time intervals).
Regarding claim 5:
Gu and Lopez teach the method of claim 2, Gu further teaches
determining a performance metric for a first machine learning model (Gu, page 1, section 1.1, “Our primary contributions are two-fold. First, we provide a set of benchmarks for the predictive accuracy of machine learning methods ... The first is a comparatively high out-of-sample predictive R2 relative to preceding literature”. Examiner notes that the R2 model is being mapped to the performance metric).
in response to determining whether the performance metric is below a threshold, modifying a set of hyperparameters of the first machine learning model (Gu, page 21, section 2.7, “Next, “early stopping” is a general machine learning regularization tool. It begins from an initial parameter guess that imposes parsimonious parameterization (for example, setting all θ values close to zero). In each step of the optimization algorithm, the parameter guesses are gradually updated to reduce prediction errors in the training sample. At each new guess, predictions are also constructed for the validation sample, and the optimization is terminated when the validation sample errors begin to increase”).
Regarding claim 6:
Gu and Lopez teach the method of claim 2, Gu further teaches
receiving new panel data and determining a new variable to uniquely identify records along a cross-sectional dimension in the new panel data (Gu, page 64, Section D Sample Splitting, “A common alternative to the fixed split scheme is a “rolling” scheme, in which the training and validation samples gradually shift forward in time to include more recent data, but holds the total number of time periods in each training and validation sample fixed. For each rolling window, one re f its the model from the prevailing training and validation samples, and tracks a model’s performance in the remaining test data that has not been subsumed by the rolling windows” and “In its most general form, we describe an asset’s excess return as an additive prediction error model: ri,t+1 = Et(ri,t+1) + i,t+1, where Et(ri,t+1) = g (zi,t). (1) (2) Stocks are indexed as i = 1,...,Nt and months by t = 1,...,T. For ease of presentation, we assume a balanced panel of stocks, and defer the discussion on missing data to Section 3.1” (Gu, page 8, section 2). Examiner notes that the shifting in time means new data is being received).
splitting, based on the new variable, the new panel data into a new train dataset, a new test dataset, and a new validation dataset (Gu, page 4, section 1.5, “We conduct a large scale empirical analysis, investigating nearly 30,000 individual stocks over 60 years from 1957 to 2016. Our predictor set includes 94 characteristics for each stock, interactions of each characteristic with eight aggregate time series variables, and 74 industry sector dummy variables, totaling more than 900 baseline signals” and “we divide our sample into three disjoint time periods that maintain the temporal ordering of the data. The first, or “training,” subsample is used to estimate the model subject to a specific set of tuning parameter values. The second, or “validation,” sample is used for tuning the hyperparameters … the third, or “testing,” subsample, which is used for neither estimation nor tuning, is truly out-of-sample and thus is used to evaluate a method’s predictive performance” (Gu, page 9, section 2.1). Examiner notes that the multiple stocks being observed over a period of time is mapped to panel data).
determining, based on the new validation dataset, one or more time intervals and removing, from the new validation dataset, data falling outside of the one or more time intervals, wherein the new validation dataset includes data that is out-of-sample and out-of-time with respect to the new train dataset (Gu , page 4, section 1.5, “The second, or “validation,” sample is used for tuning the hyperparameters”).
training a second machine learning model based on the new train dataset (Gu, page 9, section 2.1, “We follow the most common approach in the literature and select tuning parameters adaptively from the data in a validation sample. In particular, we divide our sample into three disjoint time periods that maintain the temporal ordering of the data. The first, or “training,” subsample is used to estimate the model subject to a specific set of tuning parameter values).
Gu does not teach, but Lopez does teach
removing, from the validation dataset and after splitting the panel data into the train dataset, the test dataset, and the validation dataset, data falling within an out- of-time period that is after a particular period of time of the panel data (Lopez, Chapter 7.4.2, “The embargo does not need to affect training observations prior to a test set, because training labels Yi = f[[ti, 0, ti, 1]], where ti, 1 < tj, 0 (training ends before testing begins), contain information that was available at the testing time tj, 0. In other words, we are only concerned with training labels Yi = f[[ti, 0, ti, 1]] that take place immediately after the test, tj, 1 ≤ ti, 0 ≤ tj, 1 + h.” and “In the previous sections we have discussed how to produce training/testing splits when labels overlap. That introduced the notion of purging and embargoing, in the particular context of model development. In general, we need to purge and embargo overlapping training observations whenever we produce a train/test split, whether it is for hyper-parameter fitting, backtesting, or performance evaluation” (Lopez, Chapter 7.4.3)
PNG
media_image1.png
311
633
media_image1.png
Greyscale
Examiner notes that we can apply this embargo function to the validation dataset mentioned above).
removing, from the train dataset, data falling within the out-of-time period (Lopez, Chapter 7.4.2, “The embargo does not need to affect training observations prior to a test set, because training labels Yi = f[[ti, 0, ti, 1]], where ti, 1 < tj, 0 (training ends before testing begins), contain information that was available at the testing time tj, 0. In other words, we are only concerned with training labels Yi = f[[ti, 0, ti, 1]] that take place immediately after the test, tj, 1 ≤ ti, 0 ≤ tj, 1 + h.”)
PNG
media_image1.png
311
633
media_image1.png
Greyscale
determining, based on the new validation dataset, one or more time intervals and removing, from the new validation dataset, data falling outside of the one or more time intervals, wherein the new validation dataset includes data that is out-of-sample and out-of-time with respect to the new train dataset (Lopez, chapter 7.4, “One way to reduce leakage is to purge from the training set all observations whose labels overlapped in time with those labels included in the testing set. I call this process “purging”.
PNG
media_image2.png
304
630
media_image2.png
Greyscale
Examiner notes that this function drops (removes) training data that fall within a time interval and we can apply this purging function to the validation set mentioned above).
removing, from the new train dataset, data falling within the one or more time intervals, wherein the new train dataset includes data that is out-of-sample and out-of-time with respect to the new validation dataset (Lopez, chapter 7.4, “One way to reduce leakage is to purge from the training set all observations whose labels overlapped in time with those labels included in the testing set. I call this process “purging”.
PNG
media_image2.png
304
630
media_image2.png
Greyscale
Examiner notes that this function drops (removes) training data that fall within a time interval).
Gu and Lopez are considered analogous to the claimed invention because they both deal with time-series data and cross-sectional data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gu to use Lopez’s purging and embargo function to filter data out from Gu’s training and validation datasets depending on the out-of-time period or time interval(s). One of the ordinary skill in the art would have known to apply the known technique of using the purging and embargo functions to remove data based on one or more time periods. Therefore, applying Lopez’s technique would yield the predictable result of keeping datasets out-of-time from each other to prevent data leakage (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results.
Regarding claim 7:
Gu and Lopez teach the method of claim 2, Gu further teaches
splitting the received panel data into a train dataset, a test dataset, and a validation dataset further comprises preprocessing the received panel data to remove any instances of incomplete data (Gu, page 53, Section A Monte Carlo Simulations, “We simulate the panel of characteristics for each 1 ≤ i ≤ N and each 1 ≤ j ≤ Pc from the following model: cij,t = 2 N +1CSrank(¯cij,t) − 1, ¯cij,t = ρj¯cij,t−1 + ij,t, (A.1) where ρj ∼ U[0.9,1], ij,t ∼ N(0,1 − ρ2 j) and CSrank is Cross-Section rank function, so that the characteristics feature some degree of persistence over time, yet is cross-sectionally normalized to be within [−1,1]. This matches our data cleaning procedure in the empirical study”. Examiner notes that Gu’s data cleaning procedure normalizes the data, which includes removing incomplete data).
Regarding claim 8:
Gu and Lopez teach the method of claim 2, Gu further teaches
determining, from received panel data, a variable to uniquely identify records along a cross-sectional dimension in the received panel data (Gu, page 8, section 2, “In its most general form, we describe an asset’s excess return as an additive prediction error model: ri,t+1 = Et(ri,t+1) + i,t+1, where Et(ri,t+1) = g (zi,t). (1) (2) Stocks are indexed as i = 1,...,Nt and months by t = 1,...,T. For ease of presentation, we assume a balanced panel of stocks, and defer the discussion on missing data to Section 3.1”. Examiner notes that the multiple stocks being observed over a period of time is mapped to panel data and the stocks being indexed as i shows indicates that i is the unique variable).
Regarding claim 9:
Gu and Lopez teach the product of claim 6, Gu further teaches
comparing a first machine learning model to the second machine learning model based on performance metrics to determine which machine learning model performs better (Gu, page 26, table 1)
PNG
media_image3.png
543
813
media_image3.png
Greyscale
Examiner notes that the table shows the comparison based on the R2 model.
Gu does not teach, but Lopez does teach
in response to determining that the second machine learning model outperforms the first machine learning model, averaging outputs from the first machine learning model and the second machine learning model to generate a new output (Lopez, Chapter 6.3, “Bootstrap aggregation (bagging) is an effective way of reducing the variance in forecasts. It works as follows: First, generate N training datasets by random sampling with replacement. Second, fit N estimators, one on each training set. These estimators are fit independently from each other, hence the models can be fit in parallel. Third, the ensemble forecast is the simple average of the individual forecasts from the N models”).
Gu and Lopez are considered analogous to the claimed invention because they both deal with time-series data and cross-sectional data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Gu to use Lopez’s ensemble forecast to find the average between the machine learning models from Gu. One of the ordinary skill in the art would have known to apply the known technique of averaging outputs of machine learning models. Therefore, applying Lopez’s technique would yield the predictable result of averaging two machine learning outputs (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results.
Regarding claim 10:
Gu and Lopez teach the method of claim 2, Gu further teaches
in response to determining that the performance metric is below the threshold, determining a new size for the train dataset, the test dataset, and the validation dataset, wherein the new size corresponds to a larger time period of the panel data compared to a previous size (Gu, page 64, Section D Sample Splitting, “The third is a “recursive” performance evaluation scheme. Like the rolling approach, it gradually includes more recent observations in the training and validation windows. But the recursive scheme always retains the entire history in the training sample, thus its window size gradually increases”).
Regarding claim 11:
Gu and Lopez teach the method of claim 2, Gu further teaches
the test dataset is used to generate an estimate of model error that is out-of-sample and out-of-time with respect to the validation dataset and the train dataset (Gu, page 25, section 3.1, “We divide the 60 years of data into 18 years of training sample (1957- 1974), 12 years of validation sample (1975- 1986), and the remaining 30 years (1987- 2016) for out-of-sample testing” and “Table 1 presents the comparison of machine learning techniques in terms of their out-of-sample predictive R2” (Gu, page 25, section 3.2). Examiner notes that the training sample is out-of-time compared to the validation sample and training sample).
Response to Arguments
On page 1, the Applicant argues:
The Specification is objected to because paragraphs 0010-0011 and paragraphs 0032- 0033 allegedly recite examples that "do not match with each other." Applicant respectfully submits that there is no requirement that examples in the specification must match each other. Different examples may illustrate different implementations of the claimed invention. Accordingly, Applicant respectfully requests reconsideration and withdrawal of the objection to the Specification.
Regarding the Applicant’s argument that the examples in the specification do not have to match each other, the Examiner agrees. Therefore the objection to the specification has been withdrawn.
On page 2, Applicant argues:
Claims 6 and 16 stand rejected under 35 U.S.C. § 112(b).
Without acquiescing in this rejection, but merely to expedite prosecution, Applicant amends claims 6 and 16 to address the alleged concerns.
Accordingly, Applicant respectfully requests that the Examiner reconsider and withdraw this rejection
Regarding the Applicant’s amendment to the 35 U.S.C. § 112(b), the Examiner agrees. Therefore, the 35 U.S.C. § 112(b) rejection has been withdrawn.
On page 2-3, Applicant argues:
Claims 1-20 are rejected under 35 U.S.C. § 101 as allegedly being directed to non- statutory subject matter.
For at least the reasons discussed during the interview and without acquiescing in the rejection, the amended independent claims, and the claims that depend thereon, are patent- eligible under 35 U.S.C. § 101.
Even assuming the amended independent claims can reasonably be determined to recite an abstract idea under Prong One of Step 2A, which Applicant does not concede is the case for this application, MPEP 2106.04(d)(I) states "[a] claim reciting a judicial exception is not directed to the judicial exception if it also recites additional elements demonstrating that the claim as a whole integrates the exception into a practical application. One way to demonstrate such integration is when the claimed invention improves the functioning of a computer or improves another technology or technical field" under Prong Two of Step 2A of the subject matter eligibility test.
As stated in the USPTO's December 4, 2025 Memorandum regarding the In re
Desjardins decision: "rejection was vacated because the claims were directed to training a machine learning model on multiple tasks, while preserving prior tasks performed, which- 12
properly integrated an otherwise abstract idea into a practical application; therefore, the claim satisfied Step 2 of the framework set forth in Alice Corp. v. CLS Bank Int'l, 573 U.S. 208 (2014). We specifically credited the claims for improving the functioning of the machine learning model itself."
Similarly, representative amended claim 1 as a whole integrates into a practical
application because it is directed to improvements in the technical field of machine learning. Specifically, the claimed invention provides better-performing machine learning models by ensuring the validation data contain both out-of-sample and out-of-time observations. See, for example, paragraphs 3-5 of the specification.
Since the independent claims provide at least some the technical improvements discussed above, the claims integrate any alleged abstract idea into a practical application under Prong Two of Step 2A and are therefore patent-eligible under 35 U.S.C. § 101.
Accordingly, Applicant respectfully requests that the Examiner reconsider and withdraw this rejection
Regarding the Applicant’s argument that the claimed invention is an improvement to a technology or technical field, the Examiner agrees. Therefore, the 35 U.S.C §101 rejection has been withdrawn
On page 3-4, Applicant argues:
Claims 1 and 12-20 are rejected under 35 U.S.C. § 103 as allegedly being unpatentable over GU (Empirical Asset Pricing via Machine Learning), SCHWIEP (US Patent Application Publication No. 2022/0292308), and LOPEZ (Advances in Financial Machine Learning), and claims 2-11 are rejected under 35 U.S.C. § 103 as allegedly being unpatentable over GU and LOPEZ.
For at least the reasons discussed during the interview and without acquiescing in the rejection, the cited sections of the applied references, whether taken alone or in any reasonable combination, do not disclose one or more features recited in the amended independent claims.
Therefore, the amended independent claims and the claims that depend thereon are patentable over the cited sections of the applied references, whether taken alone or in any reasonable combination.
Accordingly, Applicant respectfully requests that the Examiner reconsider and withdraw this rejection.
Regarding the Applicant’s argument that the cited combination of references fail to disclose or suggest the amended claims, the Examiner respectfully disagrees.
In claim 1, Lopez further teaches removing, from the validation dataset and after splitting the panel data into the train dataset, the test dataset, and the validation dataset, data falling within an out- of-time period that is after a particular period of time of the panel data:
(Lopez, Chapter 7.4.2, “The embargo does not need to affect training observations prior to a test set, because training labels Yi = f[[ti, 0, ti, 1]], where ti, 1 < tj, 0 (training ends before testing begins), contain information that was available at the testing time tj, 0. In other words, we are only concerned with training labels Yi = f[[ti, 0, ti, 1]] that take place immediately after the test, tj, 1 ≤ ti, 0 ≤ tj, 1 + h.” and “In the previous sections we have discussed how to produce training/testing splits when labels overlap. That introduced the notion of purging and embargoing, in the particular context of model development. In general, we need to purge and embargo overlapping training observations whenever we produce a train/test split, whether it is for hyper-parameter fitting, backtesting, or performance evaluation” (Lopez, Chapter 7.4.3)
PNG
media_image1.png
311
633
media_image1.png
Greyscale
Lopez teaches that the purge and embargo must be done whenever a dataset is split, suggesting that the removal of data is done after the split. Examiner notes that we can apply this embargo function to the validation dataset mentioned above.
Claims 2, 6, 12, and 16 follow the same reasoning as claim 1 since claims 2 and 6 are the method and claims 12 and 16 are the non-transitory computer readable medium claims. Examiner respectfully directs the Applicant to the above 103 rejection.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Racine et al. (Consistent cross-validatory model-selection for dependent data: hv-block cross-validation) discloses cross-validation methods on time-series data. Sypniewski et al. (End-to- End Neural Networks for Speech Recognition and Classification) discloses end-to-end neural networks for speech recognition and classification and additional machine learning techniques that may be used in conjunction or separately. Meyer et al. (Improving performance of spatio-temporal machine learning models using forward feature selection and target-oriented validation) discloses target-orientated cross validation on leave-location and time-out data.
Applicant’s amendment necessitated the new ground(s) of rejection presented in this Office action. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to STEVEN VO whose telephone number is (571)272-9622. The examiner can normally be reached Monday - Friday from 7-3 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.V./ Examiner, Art Unit 2148 /MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148