Prosecution Insights
Last updated: August 18, 2026
Application No. 18/405,462

PROCESS CONTROL BASED ON SIMULTANEOUS MACHINE LEARNING OF SPATIAL AND TEMPORAL RELATIONS OF TIME SERIES DATA

Non-Final OA §101§102§103§112
Filed
Jan 05, 2024
Examiner
LU, HWEI-MIN
Art Unit
Tech Center
Assignee
International Business Machines Corporation
OA Round
1 (Non-Final)
63%
Grant Probability
Moderate
1-2
OA Rounds
3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 63% of resolved cases
63%
Career Allowance Rate
146 granted / 233 resolved
+2.7% vs TC avg
Strong +40% interview lift
Without
With
+39.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
25 currently pending
Career history
263
Total Applications
across all art units

Statute-Specific Performance

§101
9.5%
-30.5% vs TC avg
§103
50.2%
+10.2% vs TC avg
§102
11.3%
-28.7% vs TC avg
§112
29.0%
-11.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 233 resolved cases

Office Action

§101 §102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This office action is in responsive to communication(s): original application filed on 01/05/2024. Claims 1-20 are pending. Claims 1, 9, and 17 are independent. Specification The use of the term "Bluetooth" in ¶ [0028] and "Wi-Fi" in ¶¶ [0029]-[0030], which is a trade name or a mark used in commerce, has been noted in this application. The term should be accompanied by the generic terminology; furthermore the term should be capitalized wherever it appears or, where appropriate, include a proper symbol indicating use in commerce such as ™, SM , or ® following the term. Although the use of trade names and marks used in commerce (i.e., trademarks, service marks, certification marks, and collective marks) are permissible in patent applications, the proprietary nature of the marks should be respected and every effort made to prevent their use in any manner which might adversely affect their validity as commercial marks. Claim Objections Claims 3, 5, 11, and 13 are objected to because of the following informalities: in Claims 3 and 11, lines 1-2, "… wherein the predicted future values are for a next step ahead …" appears to be "… wherein the predicted future data values are for a next step ahead …"; in Claim 5, lines 2-3; and Claim 13, line 2 "… performs a depth-wise separable convolution on the different features" appears to be "… performs a depth-wise separable convolution on the corresponding different features". Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-16 and 20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 1 and 9 recite the limitation "… receiving/receive (…) one or more sets of time series data values recorded … for respective features of the system … predicting/predict … data values for the features for a future time based on receiving the set of time series data values …" in lines 2-7 and 4-9 respectively, which rendering these claims indefinite because (1) i. Claims 2-8 and 10-16 are rejected for fully incorporating the deficiency of their respective base claims. Claims 8 and 16 recite the limitation "… determining an imputation loss for the first machine learning model imputing missing data values for training data; determining a prediction loss for the second machine learning model predicting future data values related to the training data …" in lines 3-6, which rendering these claims indefinite because ". Claim 20 recites the limitation "the different features" in lines 2-3. There is insufficient antecedent basis for this limitation in the claim. Clarification is required Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-6 and 9-14 are rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract idea without significantly more. Independent Claims 1 and 9 Step 1: Claim 1 is a process claim and Claim 9 is a system claim. These claims fall within at least one of the four categories of patent eligible subject matter. Step 2A Prong 1: The claim(s) recite(s) "impute/imputing, via a first model, missing data values from the one or more sets of time series data values", " predict/predicting, via a second model, data values for the features for a future time based on receiving the set of time series data values and the imputed data values from the first model as input, wherein the first and the second models implement different functions for corresponding different features of the system", and "generating one or more commands based on the predicted data values" which can be reasonably considered as mental processes (i.e., which "can be performed in the human mind, or by a human using a pen and paper") or mathematical concepts/calculations/algorithms. Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) recite(s) additional elements/limitations of "at least one processor", "a system", "a computer system" (Claim 9), "one or more memories" (Claim 9), "receive one or more sets of time series data values recorded from operation of a system for respective features of the system", "a first machine learning model", "a second machine learning model", and "control operations for the system" which only amount to "apply it" with the use of generic computer components or insignificant extra solution activity. None of the additional elements/limitations, taken alone or in combination, integrate the abstract idea into a practical application. Step 2B: The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because (a) the additional limitation/element of "receive one or more sets of time series data values recorded from operation of a system for respective features of the system" is well-understood, routine and conventional (WURC) activity similar to "receiving or transmitting data over a network" (see MPEP 2106.05(d), "Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network)"); (b) the additional limitations/elements of "a first machine learning model" and "a second machine learning model" are also well-understood, routine and conventional (WURC) activity similar to "performing repetitive calculation" (see MPEP 2106.05(d), "Performing repetitive calculations, Flook, 437 U.S. at 594, 198 USPQ2d at 199 (recomputing or readjusting alarm limit values)"); and (c) the additional limitation/element of "control operations for the system" is also well-understood, routine and conventional (WURC) activity similar to "presenting offers and gathering statistics" (see MPEP 2106.05(d), "Presenting offers and gathering statistics, OIP Techs., 788 F.3d at 1362-63, 115 USPQ2d at 1092-93"). Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea. Claims 2 and 10 Step 1: Claim 2 is a process claim and Claim 10 is a system claim. These claims fall within at least one of the four categories of patent eligible subject matter. Step 2A Prong 1: The claim(s) does/do not further recite(s) elements/limitations which can be reasonably considered as mental processes (i.e., which "can be performed in the human mind, or by a human using a pen and paper") or mathematical concepts/algorithms/calculations. Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) further recite(s) additional element/limitation of "wherein the first machine learning model comprises a fully connected neural network" which only amount to "apply it" with the use of generic computer components or insignificant extra solution activity. None of the additional elements/limitations, taken alone or in combination, integrate the abstract idea into a practical application. Step 2B: The claim(s) does/do not include further additional elements that are sufficient to amount to significantly more than the judicial exception because the additional limitation/element of "wherein the first machine learning model comprises a fully connected neural network" is also well-understood, routine and conventional (WURC) activity similar to "performing repetitive calculation" (see MPEP 2106.05(d), "Performing repetitive calculations, Flook, 437 U.S. at 594, 198 USPQ2d at 199 (recomputing or readjusting alarm limit values)"). Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea. Claims 3 and 11 Step 1: Claim 3 is a process claim and Claim 11 is a system claim. These claims fall within at least one of the four categories of patent eligible subject matter. Step 2A Prong 1: The claim(s) further recite(s) "predict/predicting data values for a future time, wherein the predicted future values are for a next step ahead and for multiple steps ahead" which can be reasonably considered as mental processes (i.e., which "can be performed in the human mind, or by a human using a pen and paper") or mathematical concepts/calculations/algorithms. Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations. Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea. Claims 4 and 12 Step 1: Claim 4 is a process claim and Claim 12 is a system claim. These claims fall within at least one of the four categories of patent eligible subject matter. Step 2A Prong 1: The claim(s) does/do not further recite(s) elements/limitations which can be reasonably considered as mental processes (i.e., which "can be performed in the human mind, or by a human using a pen and paper") or mathematical concepts/algorithms/calculations. Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) further recite(s) additional element/limitation of "wherein the second machine learning model comprises a convolutional neural network" which only amount to "apply it" with the use of generic computer components or insignificant extra solution activity. None of the additional elements/limitations, taken alone or in combination, integrate the abstract idea into a practical application. Step 2B: The claim(s) does/do not include further additional elements that are sufficient to amount to significantly more than the judicial exception because the additional limitation/element of "wherein the second machine learning model comprises a convolutional neural network" is also well-understood, routine and conventional (WURC) activity similar to "performing repetitive calculation" (see MPEP 2106.05(d), "Performing repetitive calculations, Flook, 437 U.S. at 594, 198 USPQ2d at 199 (recomputing or readjusting alarm limit values)"). Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea. Claims 5 and 13 Step 1: Claim 5 is a process claim and Claim 13 is a system claim. These claims fall within at least one of the four categories of patent eligible subject matter. Step 2A Prong 1: The claim(s) does/do not further recite(s) elements/limitations which can be reasonably considered as mental processes (i.e., which "can be performed in the human mind, or by a human using a pen and paper") or mathematical concepts/algorithms/calculations. Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) further recite(s) additional elements/limitations of "wherein the convolutional neural network comprises linear layers" and performs a depth-wise separable convolution on the different features" which only amount to "apply it" with the use of generic computer components or insignificant extra solution activity. None of the additional elements/limitations, taken alone or in combination, integrate the abstract idea into a practical application. Step 2B: The claim(s) does/do not include further additional elements that are sufficient to amount to significantly more than the judicial exception because the additional limitations/elements of "wherein the convolutional neural network comprises linear layers" and performs a depth-wise separable convolution on the different features" are also well-understood, routine and conventional (WURC) activity similar to "performing repetitive calculation" (see MPEP 2106.05(d), "Performing repetitive calculations, Flook, 437 U.S. at 594, 198 USPQ2d at 199 (recomputing or readjusting alarm limit values)"). Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea. Claims 6 and 14 Step 1: Claim 6 is a process claim and Claim 14 is a system claim. These claims fall within at least one of the four categories of patent eligible subject matter. Step 2A Prong 1: The claim(s) does/do not further recite(s) elements/limitations which can be reasonably considered as mental processes (i.e., which "can be performed in the human mind, or by a human using a pen and paper") or mathematical concepts/algorithms/calculations. Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) further recite(s) additional element/limitation of "training the convolutional neural network to simultaneously learn the different functions for the corresponding different features" which only amount to "apply it" with the use of generic computer components or insignificant extra solution activity. None of the additional elements/limitations, taken alone or in combination, integrate the abstract idea into a practical application. Step 2B: The claim(s) does/do not include further additional elements that are sufficient to amount to significantly more than the judicial exception because the additional limitation/element of "training the convolutional neural network to simultaneously learn the different functions for the corresponding different features" is also well-understood, routine and conventional (WURC) activity similar to "performing repetitive calculation" (see MPEP 2106.05(d), "Performing repetitive calculations, Flook, 437 U.S. at 594, 198 USPQ2d at 199 (recomputing or readjusting alarm limit values)"). Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea. Claims 7 and 15 Step 1: Claim 7 is a process claim and Claim 15 is a system claim. These claims fall within at least one of the four categories of patent eligible subject matter. Step 2A Prong 1: The claim(s) does/do not further recite(s) elements/limitations which can be reasonably considered as mental processes (i.e., which "can be performed in the human mind, or by a human using a pen and paper") or mathematical concepts/algorithms/calculations. Step 2A Prong 2: This judicial exception is integrated into a practical application because the claim(s) further recite(s) additional element/limitation of "training the first and the second machine learning models simultaneously based on a combined loss" which is integrated with other judicial exception elements/limitations in a meaningful way (i.e., not just "apply it") so that an improvement of a technology indicating in the specification (e.g., ¶ [0043]: "… The time series machine learning model trains spatial and temporal machine learning models for imputation and forecasting simultaneously, and may utilize signals from forecasting tasks to improve imputation. The overall quality of forecasting via use of the so-trained time series machine learning model achieves improvement.") is reflected in the claim as a whole. Therefore, these claims include patent eligible subject matters. Claims 8 and 16 Step 1: Claim 8 is a process claim and Claim 16 is a system claim. These claims fall within at least one of the four categories of patent eligible subject matter. Step 2A Prong 1: The claim(s) further recite(s) "determining an imputation loss for the first model imputing missing data values for training data", "determining a prediction loss for the second model predicting future data values related to the training data", and "combining the imputation loss and the prediction loss to determine an overall loss" which can be reasonably considered as mental processes (i.e., which "can be performed in the human mind, or by a human using a pen and paper") or mathematical concepts/calculations/algorithms. Step 2A Prong 2: This judicial exception is integrated into a practical application because the claim(s) further recite(s) additional element/limitation of "wherein the first and the second machine learning models are trained simultaneously based on the overall loss" which is integrated with other judicial exception elements/limitations in a meaningful way (i.e., not just "apply it") so that an improvement of a technology indicating in the specification (e.g., ¶ [0043]: "… The time series machine learning model trains spatial and temporal machine learning models for imputation and forecasting simultaneously, and may utilize signals from forecasting tasks to improve imputation. The overall quality of forecasting via use of the so-trained time series machine learning model achieves improvement.") is reflected in the claim as a whole. Therefore, these claims include patent eligible subject matters. Independent Claim 17 Step 1: Claim 17 is a process claim. The claim falls within at least one of the four categories of patent eligible subject matter. Step 2A Prong 1: The claim(s) further recite(s) "determining an imputation loss for the first model imputing missing time series data values for training data", "determining a prediction loss for the second model predicting future data values related to the training data", and "combining the imputation loss and the prediction loss to determine an overall loss" which can be reasonably considered as mental processes (i.e., which "can be performed in the human mind, or by a human using a pen and paper") or mathematical concepts/calculations/algorithms. Step 2A Prong 2: This judicial exception is integrated into a practical application because the claim(s) further recite(s) additional element/limitation of "wherein the system comprises a model structure with output of the first machine learning model being input into the second machine learning model … so that the first and the second machine learning models are trained simultaneously" which is integrated with other judicial exception elements/limitations in a meaningful way (i.e., not just "apply it") so that an improvement of a technology indicating in the specification (e.g., ¶ [0043]: "… The time series machine learning model trains spatial and temporal machine learning models for imputation and forecasting simultaneously, and may utilize signals from forecasting tasks to improve imputation. The overall quality of forecasting via use of the so-trained time series machine learning model achieves improvement.") is reflected in the claim as a whole. Therefore, these claims include patent eligible subject matters. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 17-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Kim et al. ("End-to-end Multi-task Learning of Missing Value Imputation and Forecasting in Time-Series Data", 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy, Jan 10-15, 2021, pp. 8849-8856), hereinafter Kim. Independent Claim 17 Kim discloses a computer-implemented method comprising: training a system comprising first and second machine learning models via: determining an imputation loss for the first machine learning model imputing missing time series data values for training data (Kim, Section I of Page 8849: adopt the generative adversarial network (GAN) architecture to estimate the true distribution of the observed data; adopted the input dropout method in our model to learn the true distribution of the data via destruction-and-reconstruction process; in contrast to previous state-of-the-art work that added noise to destroy the original data, we dropped out the known values and learned to reconstruct them; we empirically verified that input dropout of the original method can help imputation by modeling denoising effect; we designed a synthetic dataset with a known true distribution which can be used to observe the denoising effect of the model and to analyze the missing value imputation and downstream task prediction performances; the synthetic dataset that can be used to validate not only imputation and downstream task performances, but also denoising effect of the model; Section II of Page 8850: by adopting the RNN architecture, they took advantage of the network which can process temporal information and achieved a great success in multivariate time-series imputation; proposed RNN-based models with a decaying mechanism that can effectively handle varying time gaps between observed variables, wherein the missing inputs to the main network are imputed with temporally decayed statistics obtained from the observed variables; adopted GAN architecture that jointly train a discriminator and a generator on the multivariate time-series imputation task; the discriminator distinguishes the generator output and the real data, while the generator deceives the discriminator by making a realistic output; with this additional architecture, researchers successfully imputed the missing values that follow the distribution of the training data; Section III with FIG. 1 of Page 8850: Let X n n = 1 N and l n n = 1 N denote a multivariate timeseries dataset and the corresponding label, respectively; X = x 1 ,   x 2 , … , x n ∈ X n n = 1 N with label l ∈ l n n = 1 N is a multivariate sequence of length T, where x t ∈ R d is the t-th observation of X at timestamp s t ; denote the i-th variable of the t-th observation as x i t and it is either observed or missing data; M = m 1 ,   m 2 , … , m T is a mask matrix whose element m i t ∈ 0,1 indicates whether the given variable x i t is missing or not; m i t = 1 if x i t is observed, otherwise m i t = 0 ; moreover, introduce a dropped matrix X ~ , which is generated by additionally dropping a portion of the observed variables from X to make model learn how to impute missing values; The corresponding mask matrix for X ~ , is denoted as M ~ ; define a time gap matrix Δ = δ 1 ,   δ 2 , … , δ T that serves time interval from the last observation for each variable of X ~ , i.e., δ i t shows how long a variable x ~ i has been missing consecutively until the t-th time step; Fig. 1 shows an example of model inputs with colors indicating the status of each variable: white for observed variables, red for missing variables and blue for additionally dropped variables, where X is a time series data of length five with three features; s is a time stamp vector for X; l is the given prediction label for X, which can be either classification or regression label; X ~ shows input data after additional 30% of the observed variables is dropped; M and Δ is a mask matrix and a time gap matrix for X ~ , respectively; Section VI.A of Page 8854: the PhysioNet dataset consists of 4000 Intensive Care Unit(ICU) stay records which have 80:5% of unknown variables; the main purpose of the dataset is the development of methods for predicting the mortality rate in ICU; the dataset has 35 features except for patient information; every record in the dataset has a 48 hours multivariate time-series sequence; since the dataset has various time gap between each clinical measurement and sequence length, to compile the dataset as fixed size, we convert the period as an hour by rounding the time to the nearest hour; Section IV.A-IV.B with FIG. 2 of Pages 8850-8852: present the architecture of imputation module that aims to impute missing values within the data; take advantage of GAN, so that our generator (G) learns the distribution of the real data under the supervision of the discriminator (D); particularly, we used Wasserstein GAN, which is easier to train stably than original GAN, alleviating the problem of non-convergence and mode dropping phenomenon; additionally, both G and D are gated recurrent unit (GRU)-based RNNs, where G is a bidirectional RNN with decaying cell; Since G takes as an incomplete time series data, it should handle two kinds of input variables, the observed variables and the missing data; for the former, where m i t = 1 , G works as an auto-encoder, which learns the distribution of the given data while reconstructing the input values; on the other hand, G does not reconstruct the same input value but estimates a new value to fill in the missing input value; the concept of denoising auto-encoder is to reconstruct a complete input data from a partially destroyed one, essentially learning to generate a clean “noise-reduced input”; based on this idea, we propose the dropping function, which helps the model to predict appropriate values for missing values by generating missing values that we know the ground truth: namely, X ~ ; given X ~ , not X, it is possible for the model to learn how to reconstruct the values of the unknown variables, because now there exists ground truth imputation labels for the missing values; dropping additional observed variables guarantees that missing values without ground truth are imputed appropriately and makes G learn how to remove noise in both observed variables and missing variables; to sum up, G behaves as an auto-encoder for the observed variable and authentic generator for the other; in order to indicate where missing values are, we concatenate time-series data X ~ and the mask matrix M ~ and then feed it to the generator; inspired by masked language modeling of BERT, we also make M ~ work as a mask token when G handles the missing values; with this in mind, we invert M ~   and then feed it to G, so that concatenated masking has the value of one at the observed and zero at the missing; additionally, missing data need to be imputed temporarily; we replace missing values with the mean value of each variable; finally, the input of G is shown as Z = X ~ ; ( 1 - M ~ ) ,   where ; denotes a concatenation operator; i.e., G receives additional information to distinguish observed variables from missing data through M ~ ; L2 loss is minimized to decrease the distance between prediction output of G and the raw timeseries data; finally, the loss of G for reconstructing input data is as shown in equation (3); In addition to loss LG, r, the generator is also trained by minimizing the adversarial loss, fooling D using the imputation results; the input of D is the observations with missing data replaced by the output of G; D learns to distinguish the observed variables from the estimated values, classifying the former as real and the latter as fake; specifically, input variables are classified elementwise; that is, the output of D is a d × T matrix whose elements are [Symbol font/0xCE] [0; 1]; we designed D to distinguish input sequence elementwise, as there rarely exists a sequence in which all variables are observed in highly-missing time-series data; with well-trained D, the G learns to produce complete data that follows the distribution of real data fooling D; altogether, the adversarial loss for G and D is shown as equations (4)-(6); the prevailing method for adopting RNN-based model to handle incomplete data is decaying input variable or hidden state vector of RNN cells using the time gap matrix δ i t ; if a variable is consecutively missing for a long time, the reliability of hidden state from the previous time step becomes low; therefore, the missing pattern of a variable with respect to time should be considered and the hidden state vector of RNN should be decayed if a variable has been missing for a long while; therefore, we propose a decaying method that feeds k time gap vectors, δ t - k + 1 : t , into a fully connected neural network with two linear layers followed by nonlinear functions shown in equations (7)-(8), where W η and W γ are model parameters to learn, k is a hyperparameter that denotes the length of time gap sequence and Maxpool indicates max pooling operation; Eq. (7) and Eq. (8) include a linear transformation that regards time gap variations between different variables and different time steps, respectively; with the assumption that there are multiple patterns of temporal time gap variations, we map η t to a two-dimensional space instead of a vector space and perform max pooling operation, i.e., selecting one pattern out of various candidates; accordingly, γ t is in the same vector space with the hidden state vector of RNN; introduce input decaying that downscales the input data x ~ t with the rate of ( 1 - γ t ) to keep the scale of gates in GRU constant; as hidden state vectors are directly related to predicting the imputation result, they should contain the information of underlying true distribution of input data and maintain a constant scale for its value; however, if the decaying rate is repeatedly multiplied, the hidden state vectors keep decreasing and the scale of the gates fluctuates; this inconsistency of the scale can lead to inappropriate prediction results; for this reason, we decay x ~ t with the rate of ( 1 - γ t )   so that the sum of coefficients of x ~ t and h t always equals to one; as a result, the update functions of GRU with the decaying mechanism is as follows: h t - 1 = γ t   ⨀   h t - 1 ,   z t = [ 1 - γ t ; 0 ] ⊙ z t ; u t = σ W u h t - 1 ; z t + b u ,   r t = σ W r h t - 1 ; z t + b r ; h ~ t = t a n h W h r t ⊙ h t - 1 ; z t + b h ; h t = 1 - u t ⊙ h t - 1 + u t ⊙ h ~ t , where W u , W r and W h are model parameters to learn; in Eq. (IV-B), we pad ( 1 - γ t ) with zero vectors to have the shape of z t ; note that it results in multiplying zero matrix and the concatenated m ~ t ); determining a prediction loss for the second machine learning model predicting future data values related to the training data, wherein the system comprises a model structure with output of the first machine learning model being input into the second machine learning model; and combining the imputation loss and the prediction loss to determine an overall loss (Kim, Section I of Page 8849: introduce a novel deep learning algorithm that gates real data, as well as missing ones, to remove possible noise from our input to downstream task modules; propose a time decay mechanism that considers the sequential changes in the time interval; in addition, we discovered that the decaying of the previous hidden state vector makes the current output hidden state vector small; we address this problem with the supplementation of the hidden state vector with the current input vector; Section III.C in Page 8852-8853 with FIG. 2 in Page 8851 and FIG. 3 in Page 8853: introduce a novel deep learning model using the gating mechanism, gated classifier C, which considers the reliability of imputation results and substitutes the results partially with other values based on the reliability; it is proposed based on the question of suitability of imputation results based on auto-encoder or generative adversarial training for predicting downstream task; if the missing rate of input data is very high, there is little information that we can extract to estimate its distribution; if so, the output of the imputation module might not be reliable; since our prediction label is highly related to the true distribution, unreliable noisy outputs should not be used for downstream task; therefore, we propose gated classifier that mixes observations and predictions of the imputation module based on the reliability of each value and uses the resulting data for classification; gated classifier includes gating module which estimates the suitability of the output of the generator compared to the raw input data for predicting downstream task; Fig. 3 shows how a gating matrix is introduced; first, we define a gating matrix Λ ∈ R d × T that denotes the degree of reliability for Y, the output of the generator; the gating value, λ t i ∈ 0 ,   1 is an element of Λ   that quantifies the reliability of corresponding variable y t i ; we concatenate X ~ and Y to estimate relative confidence of Y compared to X ~ ; since there exists missing values substituted with the mean value in Y, we also concatenate inverted M ~ to indicate the location of missing values similarly to the formulation of generator’s input; note that the concatenation result X ~ T ; Y T ; 1 - M ~ T T is a 3d×T matrix; next, we compute the convolution of the concatenation with various kernel sizes; regardless of the kernel size, each convolution is a mapping R 3 d × T → R d × T , zero padding the input if necessary; i.e., the number of its output channels always equals the number of variables of X ~ ; each kernel sees the entire variable of raw data, generator output and inverted mask matrix at time t at once, so that gating module can make a decision about t-th variable based on the full information at time t; . As a consequence, the kernel size is (3d, *), where d decides how much temporal information is available; after convolutions, we conduct a max-pooling operation on convolution filter outputs so as to select one logit for each variable; as a result, our gating module output has the same dimension with raw data; finally, we apply sigmoid function to the output and it is used as a gating matrix; after we obtain the gating matrix, Λ , we mix the generator output and raw data by the ratio of gating value S = Y ⊙ Λ + X ~ ⊙ 1 - Λ , where S indicates the mixed output; it is what the gating module decides to be the best input data for performing downstream task; however, since we replaced missing values in X ~ with the mean value, which is zero due to the normalization, the mixing operation causes a difference in the scale of observations and the missing values where m ~ t i = 1 ; therefore, we compensate the shortfall with GRU cell weights; it should be filled by a ratio of (1 - λ t i ) for each missing values and no compensation is necessary for observed variables; thus, we multiply (1 - λ t i ) and (1 - M ~ ) to indicates the deficiency ratio to fill-in for each variable; and then, it is concatenated with S and fed to GRU; consequently, the first-layer weights of GRU are multiplied with the ratios and S with the former compensating the latter: W c s t ; λ t ∙ 1 - m ~ t ; the GRU layer of gated classifier is then followed by a fully-connected layer that outputs the classification results; the gated classifier and the imputation module, G and D, are jointly trained in an end-to-end manner; accordingly, with the assumption that the prediction label is strongly related to true distribution of the data, G generates better results that follow true distribution while minimizing classification loss) so that the first and the second machine learning models are trained simultaneously (Kim, Section I of Page 8849: by jointly training the downstream task module and gating mechanism with adversarial loss, our model produces realistic and helpful imputation to predict the downstream task; a novel end-to-end model that imputes missing values and performs downstream tasks simultaneously with gating module and input dropping; the achievement of state-of-the- art performance on a real-world dataset; Section II of Page 8850: Cao et al. (2018) improved the downstream task prediction performance by training the imputation and the downstream task simultaneously, which greatly helped the model to generate appropriate imputation results which can also make high quality predictions for the downstream task; adopted GAN architecture that jointly train a discriminator and a generator on the multivariate time-series imputation task; Section III of Page 8850: addresses the problem of solving downstream tasks and imputing unknown variables within the given data jointly; in other words, predicting l and X using X ~ , M ~ and Δ ; propose an end-to-end GANs-based model that performs missing value imputation and downstream classification jointly; Section IV of Page 8850 with FIG. 2 in Page 8851: model is designed for end-to-end missing value imputation and downstream task prediction in multivariate time-series data; the model consists of two modules: the imputation and the prediction module; Fig. 2 shows an overview of the proposed model, wherein y t and p t denotes the prediction output of the generator and the discriminator at time t, respectively; l ^ indicates the prediction output of the gated classifier for downstream task; the proposed decaying mechanism explained is adopted to the generator with its cells receiving δ t ; on the other hand, the discriminator and the gated classifier is GRU without any decaying). Claim 18 Kim discloses all the elements as stated in Claim 17 and further discloses wherein the first machine learning model comprises a fully connected neural network (Kim, Section IV.A-IV.B with FIG. 2 of Pages 8850-8852: present the architecture of imputation module that aims to impute missing values within the data; take advantage of GAN, so that our generator (G) learns the distribution of the real data under the supervision of the discriminator (D); particularly, we used Wasserstein GAN, which is easier to train stably than original GAN, alleviating the problem of non-convergence and mode dropping phenomenon; additionally, both G and D are gated recurrent unit (GRU)-based RNNs, where G is a bidirectional RNN with decaying cell; Since G takes as an incomplete time series data, it should handle two kinds of input variables, the observed variables and the missing data; for the former, where m i t = 1 , G works as an auto-encoder, which learns the distribution of the given data while reconstructing the input values; on the other hand, G does not reconstruct the same input value but estimates a new value to fill in the missing input value; the concept of denoising auto-encoder is to reconstruct a complete input data from a partially destroyed one, essentially learning to generate a clean “noise-reduced input”; based on this idea, we propose the dropping function, which helps the model to predict appropriate values for missing values by generating missing values that we know the ground truth: namely, X ~ ; given X ~ , not X, it is possible for the model to learn how to reconstruct the values of the unknown variables, because now there exists ground truth imputation labels for the missing values; dropping additional observed variables guarantees that missing values without ground truth are imputed appropriately and makes G learn how to remove noise in both observed variables and missing variables; to sum up, G behaves as an auto-encoder for the observed variable and authentic generator for the other; in order to indicate where missing values are, we concatenate time-series data X ~ and the mask matrix M ~ and then feed it to the generator; inspired by masked language modeling of BERT, we also make M ~ work as a mask token when G handles the missing values; with this in mind, we invert M ~   and then feed it to G, so that concatenated masking has the value of one at the observed and zero at the missing; additionally, missing data need to be imputed temporarily; we replace missing values with the mean value of each variable; finally, the input of G is shown as Z = X ~ ; ( 1 - M ~ ) ,   where ; denotes a concatenation operator; i.e., G receives additional information to distinguish observed variables from missing data through M ~ ; L2 loss is minimized to decrease the distance between prediction output of G and the raw timeseries data; finally, the loss of G for reconstructing input data is as shown in equation (3); In addition to loss LG, r, the generator is also trained by minimizing the adversarial loss, fooling D using the imputation results; the input of D is the observations with missing data replaced by the output of G; D learns to distinguish the observed variables from the estimated values, classifying the former as real and the latter as fake; specifically, input variables are classified elementwise; that is, the output of D is a d × T matrix whose elements are [Symbol font/0xCE] [0; 1]; we designed D to distinguish input sequence elementwise, as there rarely exists a sequence in which all variables are observed in highly-missing time-series data; with well-trained D, the G learns to produce complete data that follows the distribution of real data fooling D; altogether, the adversarial loss for G and D is shown as equations (4)-(6); the prevailing method for adopting RNN-based model to handle incomplete data is decaying input variable or hidden state vector of RNN cells using the time gap matrix δ i t ; if a variable is consecutively missing for a long time, the reliability of hidden state from the previous time step becomes low; therefore, the missing pattern of a variable with respect to time should be considered and the hidden state vector of RNN should be decayed if a variable has been missing for a long while; therefore, we propose a decaying method that feeds k time gap vectors, δ t - k + 1 : t , into a fully connected neural network with two linear layers followed by nonlinear functions shown in equations (7)-(8), where W η and W γ are model parameters to learn, k is a hyperparameter that denotes the length of time gap sequence and Maxpool indicates max pooling operation; Eq. (7) and Eq. (8) include a linear transformation that regards time gap variations between different variables and different time steps, respectively; with the assumption that there are multiple patterns of temporal time gap variations, we map η t to a two-dimensional space instead of a vector space and perform max pooling operation, i.e., selecting one pattern out of various candidates; accordingly, γ t is in the same vector space with the hidden state vector of RNN; introduce input decaying that downscales the input data x ~ t with the rate of ( 1 - γ t ) to keep the scale of gates in GRU constant; as hidden state vectors are directly related to predicting the imputation result, they should contain the information of underlying true distribution of input data and maintain a constant scale for its value; however, if the decaying rate is repeatedly multiplied, the hidden state vectors keep decreasing and the scale of the gates fluctuates; this inconsistency of the scale can lead to inappropriate prediction results; for this reason, we decay x ~ t with the rate of ( 1 - γ t )   so that the sum of coefficients of x ~ t and h t always equals to one; as a result, the update functions of GRU with the decaying mechanism is as follows: h t - 1 = γ t   ⨀   h t - 1 ,   z t = [ 1 - γ t ; 0 ] ⊙ z t ; u t = σ W u h t - 1 ; z t + b u ,   r t = σ W r h t - 1 ; z t + b r ; h ~ t = t a n h W h r t ⊙ h t - 1 ; z t + b h ; h t = 1 - u t ⊙ h t - 1 + u t ⊙ h ~ t , where W u , W r and W h are model parameters to learn; in Eq. (IV-B), we pad ( 1 - γ t ) with zero vectors to have the shape of z t ; note that it results in multiplying zero matrix and the concatenated m ~ t ). Claim 19 Kim discloses all the elements as stated in Claim 17 and further discloses wherein the second machine learning model comprises a convolutional neural network (Kim, Section III.C in Page 8852-8853 with FIG. 2 in Page 8851 and FIG. 3 in Page 8853: introduce a novel deep learning model using the gating mechanism, gated classifier C, which considers the reliability of imputation results and substitutes the results partially with other values based on the reliability; it is proposed based on the question of suitability of imputation results based on auto-encoder or generative adversarial training for predicting downstream task; if the missing rate of input data is very high, there is little information that we can extract to estimate its distribution; if so, the output of the imputation module might not be reliable; since our prediction label is highly related to the true distribution, unreliable noisy outputs should not be used for downstream task; therefore, we propose gated classifier that mixes observations and predictions of the imputation module based on the reliability of each value and uses the resulting data for classification; gated classifier includes gating module which estimates the suitability of the output of the generator compared to the raw input data for predicting downstream task; Fig. 3 shows how a gating matrix is introduced; first, we define a gating matrix Λ ∈ R d × T that denotes the degree of reliability for Y, the output of the generator; the gating value, λ t i ∈ 0 ,   1 is an element of Λ   that quantifies the reliability of corresponding variable y t i ; we concatenate X ~ and Y to estimate relative confidence of Y compared to X ~ ; since there exists missing values substituted with the mean value in Y, we also concatenate inverted M ~ to indicate the location of missing values similarly to the formulation of generator’s input; note that the concatenation result X ~ T ; Y T ; 1 - M ~ T T is a 3d×T matrix; next, we compute the convolution of the concatenation with various kernel sizes; regardless of the kernel size, each convolution is a mapping R 3 d × T → R d × T , zero padding the input if necessary; i.e., the number of its output channels always equals the number of variables of X ~ ; each kernel sees the entire variable of raw data, generator output and inverted mask matrix at time t at once, so that gating module can make a decision about t-th variable based on the full information at time t; . As a consequence, the kernel size is (3d, *), where d decides how much temporal information is available; after convolutions, we conduct a max-pooling operation on convolution filter outputs so as to select one logit for each variable; as a result, our gating module output has the same dimension with raw data; finally, we apply sigmoid function to the output and it is used as a gating matrix; after we obtain the gating matrix, Λ , we mix the generator output and raw data by the ratio of gating value S = Y ⊙ Λ + X ~ ⊙ 1 - Λ , where S indicates the mixed output; it is what the gating module decides to be the best input data for performing downstream task; however, since we replaced missing values in X ~ with the mean value, which is zero due to the normalization, the mixing operation causes a difference in the scale of observations and the missing values where m ~ t i = 1 ; therefore, we compensate the shortfall with GRU cell weights; it should be filled by a ratio of (1 - λ t i ) for each missing values and no compensation is necessary for observed variables; thus, we multiply (1 - λ t i ) and (1 - M ~ ) to indicates the deficiency ratio to fill-in for each variable; and then, it is concatenated with S and fed to GRU; consequently, the first-layer weights of GRU are multiplied with the ratios and S with the former compensating the latter: W c s t ; λ t ∙ 1 - m ~ t ; the GRU layer of gated classifier is then followed by a fully-connected layer that outputs the classification results; the gated classifier and the imputation module, G and D, are jointly trained in an end-to-end manner; accordingly, with the assumption that the prediction label is strongly related to true distribution of the data, G generates better results that follow true distribution while minimizing classification loss). Claim 20 Kim discloses all the elements as stated in Claim 19 and further discloses wherein the convolutional neural network comprises linear layers and performs a depth-wise separable convolution on the different features (Kim, Section III.C in Page 8852-8853 with FIG. 2 in Page 8851 and FIG. 3 in Page 8853: introduce a novel deep learning model using the gating mechanism, gated classifier C, which considers the reliability of imputation results and substitutes the results partially with other values based on the reliability; it is proposed based on the question of suitability of imputation results based on auto-encoder or generative adversarial training for predicting downstream task; if the missing rate of input data is very high, there is little information that we can extract to estimate its distribution; if so, the output of the imputation module might not be reliable; since our prediction label is highly related to the true distribution, unreliable noisy outputs should not be used for downstream task; therefore, we propose gated classifier that mixes observations and predictions of the imputation module based on the reliability of each value and uses the resulting data for classification; gated classifier includes gating module which estimates the suitability of the output of the generator compared to the raw input data for predicting downstream task; Fig. 3 shows how a gating matrix is introduced; first, we define a gating matrix Λ ∈ R d × T that denotes the degree of reliability for Y, the output of the generator; the gating value, λ t i ∈ 0 ,   1 is an element of Λ   that quantifies the reliability of corresponding variable y t i ; we concatenate X ~ and Y to estimate relative confidence of Y compared to X ~ ; since there exists missing values substituted with the mean value in Y, we also concatenate inverted M ~ to indicate the location of missing values similarly to the formulation of generator’s input; note that the concatenation result X ~ T ; Y T ; 1 - M ~ T T is a 3d×T matrix; next, we compute the convolution of the concatenation with various kernel sizes; regardless of the kernel size, each convolution is a mapping R 3 d × T → R d × T , zero padding the input if necessary; i.e., the number of its output channels always equals the number of variables of X ~ ; each kernel sees the entire variable of raw data, generator output and inverted mask matrix at time t at once, so that gating module can make a decision about t-th variable based on the full information at time t; . As a consequence, the kernel size is (3d, *), where d decides how much temporal information is available; after convolutions, we conduct a max-pooling operation on convolution filter outputs so as to select one logit for each variable; as a result, our gating module output has the same dimension with raw data; finally, we apply sigmoid function to the output and it is used as a gating matrix; after we obtain the gating matrix, Λ , we mix the generator output and raw data by the ratio of gating value S = Y ⊙ Λ + X ~ ⊙ 1 - Λ , where S indicates the mixed output; it is what the gating module decides to be the best input data for performing downstream task; however, since we replaced missing values in X ~ with the mean value, which is zero due to the normalization, the mixing operation causes a difference in the scale of observations and the missing values where m ~ t i = 1 ; therefore, we compensate the shortfall with GRU cell weights; it should be filled by a ratio of (1 - λ t i ) for each missing values and no compensation is necessary for observed variables; thus, we multiply (1 - λ t i ) and (1 - M ~ ) to indicates the deficiency ratio to fill-in for each variable; and then, it is concatenated with S and fed to GRU; consequently, the first-layer weights of GRU are multiplied with the ratios and S with the former compensating the latter: W c s t ; λ t ∙ 1 - m ~ t ; the GRU layer of gated classifier is then followed by a fully-connected layer that outputs the classification results; the gated classifier and the imputation module, G and D, are jointly trained in an end-to-end manner; accordingly, with the assumption that the prediction label is strongly related to true distribution of the data, G generates better results that follow true distribution while minimizing classification loss). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-16 are rejected under 35 U.S.C. 103 as being unpatentable over Kim in view of Ren et al. ("Incremental Bayesian tensor learning for structural monitoring data imputation and response forecasting", arXiv:2007.00790v3, Jul 17, 2020, pp. 1-25), hereinafter Ren. Independent Claims 1 and 9 Kim discloses a computer-implemented method comprising: receiving, via at least one processor, one or more sets of time series data values recorded from operation of a system for respective features of the system (Kim, ABSTRACT of Page 8849: multivariate time-series prediction is a common task, but it often becomes challenging due to missing data caused by unreliable sensors and other issues; in fact, inaccurate imputation of missing values can degrade the downstream prediction performance, so it may be better not to rely on the estimated values of missing data; furthermore, observed data may contain noise, so denoising them can be helpful for the main task at hand; Section I of Page 8849: multivariate time-series analysis is applied in numerous research areas such as weather forecasting, traffic analysis, patient diagnosis, and economics; such data are often partially observed, limiting their applicability in real-world applications; missing values are not only caused by the hardware problems and human error during data collection but also due to the difficulty in reliably acquiring data for, say, clinical sensor data; even with such partially missing, incomplete data, one still has to perform downstream tasks such as predicting the patient mortality rate in an intensive care unit (ICU) [2] or forecasting financial time series; Section III with FIG. 1 of Page 8850: Let X n n = 1 N and l n n = 1 N denote a multivariate timeseries dataset and the corresponding label, respectively; X = x 1 ,   x 2 , … , x n ∈ X n n = 1 N with label l ∈ l n n = 1 N is a multivariate sequence of length T, where x t ∈ R d is the t-th observation of X at timestamp s t ; denote the i-th variable of the t-th observation as x i t and it is either observed or missing data; M = m 1 ,   m 2 , … , m T is a mask matrix whose element m i t ∈ 0,1 indicates whether the given variable x i t is missing or not; m i t = 1 if x i t is observed, otherwise m i t = 0 ; moreover, introduce a dropped matrix X ~ , which is generated by additionally dropping a portion of the observed variables from X to make model learn how to impute missing values; The corresponding mask matrix for X ~ , is denoted as M ~ ; define a time gap matrix Δ = δ 1 ,   δ 2 , … , δ T that serves time interval from the last observation for each variable of X ~ , i.e., δ i t shows how long a variable x ~ i has been missing consecutively until the t-th time step; Fig. 1 shows an example of model inputs with colors indicating the status of each variable: white for observed variables, red for missing variables and blue for additionally dropped variables, where X is a time series data of length five with three features; s is a time stamp vector for X; l is the given prediction label for X, which can be either classification or regression label; X ~ shows input data after additional 30% of the observed variables is dropped; M and Δ is a mask matrix and a time gap matrix for X ~ , respectively; Section VI.A of Page 8854: the PhysioNet dataset consists of 4000 Intensive Care Unit(ICU) stay records which have 80:5% of unknown variables; the main purpose of the dataset is the development of methods for predicting the mortality rate in ICU; the dataset has 35 features except for patient information; every record in the dataset has a 48 hours multivariate time-series sequence; since the dataset has various time gap between each clinical measurement and sequence length, to compile the dataset as fixed size, we convert the period as an hour by rounding the time to the nearest hour); imputing, via a first machine learning model, missing data values from the one or more sets of time series data values (Kim, Section I of Page 8849: adopt the generative adversarial network (GAN) architecture to estimate the true distribution of the observed data; adopted the input dropout method in our model to learn the true distribution of the data via destruction-and-reconstruction process; in contrast to previous state-of-the-art work that added noise to destroy the original data, we dropped out the known values and learned to reconstruct them; we empirically verified that input dropout of the original method can help imputation by modeling denoising effect; we designed a synthetic dataset with a known true distribution which can be used to observe the denoising effect of the model and to analyze the missing value imputation and downstream task prediction performances; the synthetic dataset that can be used to validate not only imputation and downstream task performances, but also denoising effect of the model; Section II of Page 8850: by adopting the RNN architecture, they took advantage of the network which can process temporal information and achieved a great success in multivariate time-series imputation; proposed RNN-based models with a decaying mechanism that can effectively handle varying time gaps between observed variables, wherein the missing inputs to the main network are imputed with temporally decayed statistics obtained from the observed variables; adopted GAN architecture that jointly train a discriminator and a generator on the multivariate time-series imputation task; the discriminator distinguishes the generator output and the real data, while the generator deceives the discriminator by making a realistic output; with this additional architecture, researchers successfully imputed the missing values that follow the distribution of the training data; Section IV.A-IV.B with FIG. 2 of Pages 8850-8852: present the architecture of imputation module that aims to impute missing values within the data; take advantage of GAN, so that our generator (G) learns the distribution of the real data under the supervision of the discriminator (D); particularly, we used Wasserstein GAN, which is easier to train stably than original GAN, alleviating the problem of non-convergence and mode dropping phenomenon; additionally, both G and D are gated recurrent unit (GRU)-based RNNs, where G is a bidirectional RNN with decaying cell; Since G takes as an incomplete time series data, it should handle two kinds of input variables, the observed variables and the missing data; for the former, where m i t = 1 , G works as an auto-encoder, which learns the distribution of the given data while reconstructing the input values; on the other hand, G does not reconstruct the same input value but estimates a new value to fill in the missing input value; the concept of denoising auto-encoder is to reconstruct a complete input data from a partially destroyed one, essentially learning to generate a clean “noise-reduced input”; based on this idea, we propose the dropping function, which helps the model to predict appropriate values for missing values by generating missing values that we know the ground truth: namely, X ~ ; given X ~ , not X, it is possible for the model to learn how to reconstruct the values of the unknown variables, because now there exists ground truth imputation labels for the missing values; dropping additional observed variables guarantees that missing values without ground truth are imputed appropriately and makes G learn how to remove noise in both observed variables and missing variables; to sum up, G behaves as an auto-encoder for the observed variable and authentic generator for the other; in order to indicate where missing values are, we concatenate time-series data X ~ and the mask matrix M ~ and then feed it to the generator; inspired by masked language modeling of BERT, we also make M ~ work as a mask token when G handles the missing values; with this in mind, we invert M ~   and then feed it to G, so that concatenated masking has the value of one at the observed and zero at the missing; additionally, missing data need to be imputed temporarily; we replace missing values with the mean value of each variable; finally, the input of G is shown as Z = X ~ ; ( 1 - M ~ ) ,   where ; denotes a concatenation operator; i.e., G receives additional information to distinguish observed variables from missing data through M ~ ; L2 loss is minimized to decrease the distance between prediction output of G and the raw timeseries data; finally, the loss of G for reconstructing input data is as shown in equation (3); In addition to loss LG, r, the generator is also trained by minimizing the adversarial loss, fooling D using the imputation results; the input of D is the observations with missing data replaced by the output of G; D learns to distinguish the observed variables from the estimated values, classifying the former as real and the latter as fake; specifically, input variables are classified elementwise; that is, the output of D is a d × T matrix whose elements are [Symbol font/0xCE] [0; 1]; we designed D to distinguish input sequence elementwise, as there rarely exists a sequence in which all variables are observed in highly-missing time-series data; with well-trained D, the G learns to produce complete data that follows the distribution of real data fooling D; altogether, the adversarial loss for G and D is shown as equations (4)-(6); the prevailing method for adopting RNN-based model to handle incomplete data is decaying input variable or hidden state vector of RNN cells using the time gap matrix δ i t ; if a variable is consecutively missing for a long time, the reliability of hidden state from the previous time step becomes low; therefore, the missing pattern of a variable with respect to time should be considered and the hidden state vector of RNN should be decayed if a variable has been missing for a long while; therefore, we propose a decaying method that feeds k time gap vectors, δ t - k + 1 : t , into a fully connected neural network with two linear layers followed by nonlinear functions shown in equations (7)-(8), where W η and W γ are model parameters to learn, k is a hyperparameter that denotes the length of time gap sequence and Maxpool indicates max pooling operation; Eq. (7) and Eq. (8) include a linear transformation that regards time gap variations between different variables and different time steps, respectively; with the assumption that there are multiple patterns of temporal time gap variations, we map η t to a two-dimensional space instead of a vector space and perform max pooling operation, i.e., selecting one pattern out of various candidates; accordingly, γ t is in the same vector space with the hidden state vector of RNN; introduce input decaying that downscales the input data x ~ t with the rate of ( 1 - γ t ) to keep the scale of gates in GRU constant; as hidden state vectors are directly related to predicting the imputation result, they should contain the information of underlying true distribution of input data and maintain a constant scale for its value; however, if the decaying rate is repeatedly multiplied, the hidden state vectors keep decreasing and the scale of the gates fluctuates; this inconsistency of the scale can lead to inappropriate prediction results; for this reason, we decay x ~ t with the rate of ( 1 - γ t )   so that the sum of coefficients of x ~ t and h t always equals to one; as a result, the update functions of GRU with the decaying mechanism is as follows: h t - 1 = γ t   ⨀   h t - 1 ,   z t = [ 1 - γ t ; 0 ] ⊙ z t ; u t = σ W u h t - 1 ; z t + b u ,   r t = σ W r h t - 1 ; z t + b r ; h ~ t = t a n h W h r t ⊙ h t - 1 ; z t + b h ; h t = 1 - u t ⊙ h t - 1 + u t ⊙ h ~ t , where W u , W r and W h are model parameters to learn; in Eq. (IV-B), we pad ( 1 - γ t ) with zero vectors to have the shape of z t ; note that it results in multiplying zero matrix and the concatenated m ~ t ); predicting, via a second machine learning model, data values for the features for a future time based on receiving the set of time series data values and the imputed data values from the first machine learning model as input (Kim, Section I of Page 8849: introduce a novel deep learning algorithm that gates real data, as well as missing ones, to remove possible noise from our input to downstream task modules; propose a time decay mechanism that considers the sequential changes in the time interval; in addition, we discovered that the decaying of the previous hidden state vector makes the current output hidden state vector small; we address this problem with the supplementation of the hidden state vector with the current input vector; Section III.C in Page 8852-8853 with FIG. 2 in Page 8851 and FIG. 3 in Page 8853: introduce a novel deep learning model using the gating mechanism, gated classifier C, which considers the reliability of imputation results and substitutes the results partially with other values based on the reliability; it is proposed based on the question of suitability of imputation results based on auto-encoder or generative adversarial training for predicting downstream task; if the missing rate of input data is very high, there is little information that we can extract to estimate its distribution; if so, the output of the imputation module might not be reliable; since our prediction label is highly related to the true distribution, unreliable noisy outputs should not be used for downstream task; therefore, we propose gated classifier that mixes observations and predictions of the imputation module based on the reliability of each value and uses the resulting data for classification; gated classifier includes gating module which estimates the suitability of the output of the generator compared to the raw input data for predicting downstream task; Fig. 3 shows how a gating matrix is introduced; first, we define a gating matrix Λ ∈ R d × T that denotes the degree of reliability for Y, the output of the generator; the gating value, λ t i ∈ 0 ,   1 is an element of Λ   that quantifies the reliability of corresponding variable y t i ; we concatenate X ~ and Y to estimate relative confidence of Y compared to X ~ ; since there exists missing values substituted with the mean value in Y, we also concatenate inverted M ~ to indicate the location of missing values similarly to the formulation of generator’s input; note that the concatenation result X ~ T ; Y T ; 1 - M ~ T T is a 3d×T matrix; next, we compute the convolution of the concatenation with various kernel sizes; regardless of the kernel size, each convolution is a mapping R 3 d × T → R d × T , zero padding the input if necessary; i.e., the number of its output channels always equals the number of variables of X ~ ; each kernel sees the entire variable of raw data, generator output and inverted mask matrix at time t at once, so that gating module can make a decision about t-th variable based on the full information at time t; . As a consequence, the kernel size is (3d, *), where d decides how much temporal information is available; after convolutions, we conduct a max-pooling operation on convolution filter outputs so as to select one logit for each variable; as a result, our gating module output has the same dimension with raw data; finally, we apply sigmoid function to the output and it is used as a gating matrix; after we obtain the gating matrix, Λ , we mix the generator output and raw data by the ratio of gating value S = Y ⊙ Λ + X ~ ⊙ 1 - Λ , where S indicates the mixed output; it is what the gating module decides to be the best input data for performing downstream task; however, since we replaced missing values in X ~ with the mean value, which is zero due to the normalization, the mixing operation causes a difference in the scale of observations and the missing values where m ~ t i = 1 ; therefore, we compensate the shortfall with GRU cell weights; it should be filled by a ratio of (1 - λ t i ) for each missing values and no compensation is necessary for observed variables; thus, we multiply (1 - λ t i ) and (1 - M ~ ) to indicates the deficiency ratio to fill-in for each variable; and then, it is concatenated with S and fed to GRU; consequently, the first-layer weights of GRU are multiplied with the ratios and S with the former compensating the latter: W c s t ; λ t ∙ 1 - m ~ t ; the GRU layer of gated classifier is then followed by a fully-connected layer that outputs the classification results; the gated classifier and the imputation module, G and D, are jointly trained in an end-to-end manner; accordingly, with the assumption that the prediction label is strongly related to true distribution of the data, G generates better results that follow true distribution while minimizing classification loss), wherein the first and the second machine learning models implement different functions for corresponding different features of the system (Kim, Section I of Page 8849: by jointly training the downstream task module and gating mechanism with adversarial loss, our model produces realistic and helpful imputation to predict the downstream task; a novel end-to-end model that imputes missing values and performs downstream tasks simultaneously with gating module and input dropping; the achievement of state-of-the- art performance on a real-world dataset; Section II of Page 8850: Cao et al. (2018) improved the downstream task prediction performance by training the imputation and the downstream task simultaneously, which greatly helped the model to generate appropriate imputation results which can also make high quality predictions for the downstream task; adopted GAN architecture that jointly train a discriminator and a generator on the multivariate time-series imputation task; Section III of Page 8850: addresses the problem of solving downstream tasks and imputing unknown variables within the given data jointly; in other words, predicting l and X using X ~ , M ~ and Δ ; propose an end-to-end GANs-based model that performs missing value imputation and downstream classification jointly; Section IV of Page 8850 with FIG. 2 in Page 8851: model is designed for end-to-end missing value imputation and downstream task prediction in multivariate time-series data; the model consists of two modules: the imputation and the prediction module; Fig. 2 shows an overview of the proposed model, wherein y t and p t denotes the prediction output of the generator and the discriminator at time t, respectively; l ^ indicates the prediction output of the gated classifier for downstream task; the proposed decaying mechanism explained is adopted to the generator with its cells receiving δ t ; on the other hand, the discriminator and the gated classifier is GRU without any decaying); and . Kim further discloses a computer system comprising: one or more memories; and at least one processor coupled to the one or more memories, and configured to perform the method describe above (Kim, Section I of Page 8849: inherent in a system to jointly training the downstream task module and gating mechanism with adversarial loss). Kim fails to explicitly disclose generating, via the at least one processor, one or more commands for control operations for the system based on the predicted data values. Ren teaches a system and a method relating to Data Imputation and Response Forecasting (Ren, TITLE and Abstract of Page 1), wherein generating, via the at least one processor, one or more commands for control operations for the system based on the predicted data values (Ren, Abstract of Page 1: there has been increased interest in missing sensor data imputation, which is ubiquitous in the field of structural health monitoring (SHM) due to discontinuous sensing caused by sensor malfunction; recent development in Bayesian temporal factorization models for high-dimensional time series analysis has provided an effective tool solve both imputation and prediction problems; however, for large datasets, the default Bayesian temporal factorization model becomes less inefficient since the model has to be fully retrained when new data arrives; a potential solution is to train the model using a short time window covering only most recent data; however, by doing so, we may miss some critical dynamics and long-term dependencies which can only be identified from a longer time window; to address this fundamental issue in temporal factorization models, this paper presents an incremental Bayesian tensor learning scheme to achieve efficient imputation and prediction of structural response in long-term SHM; in particular, a spatiotemporal tensor is first constructed followed by Bayesian tensor factorization that extracts latent features for missing data imputation; to enable structural response forecasting based on long-term and incomplete sensing data, we develop an incremental learning scheme to effectively update the Bayesian temporal factorization model; the performance of the proposed approach is validated on continuous field-sensing data (including strain and temperature records) of a concrete bridge, based on the assumption that strain time histories are highly correlated to temperature recordings; the results indicate that the proposed probabilistic tensor learning framework is accurate and robust even in the presence of large rates of random missing, structured missing and their combination; the effect of rank selection on the imputation and prediction performance is also investigated; the results show that a better estimation accuracy can be achieved with a higher rank for random missing whereas a lower rank for structured missing; Section 1 of Pages 1-2: high-quality data plays a pivotal role in structural health monitoring (SHM) for condition assessment, damage detection, and decision making; however, during long-term monitoring, it is inevitable for imperfect and corrupted sensor measurements, especially in a harsh and noisy environment, which calls for effective approaches for imputation/recovery missing and noisy data; furthermore, in order to conduct real-time early-warning of structural deterioration or even disastrous failure, forecasting/prediction of structural response has also received considerable attention; the general idea of time series analysis, in the context of imputation and forecasting, is to find key dynamic patterns from observations and establish a mapping function between the historical records and the estimation; nevertheless, these tasks are rather challenging on account of complex spatiotemporal dependencies and inherent difficulty in large-scale and nonlinear characteristics of SHM data, especially in practical applications; i.e., conduct real-time early-warning of structural deterioration or even disastrous failure based on forecasting/prediction of structural response; Section 2 and Section 2.1 in Pages 3-4: formulate the problem of SHM data imputation and response forecasting in the context of incremental Bayesian tensor learning, and present the spatiotemporal dependency modeling procedure via matrix factorization; the goal of continuous/steaming SHM data imputation and forecasting is to estimate the missing values and predict the future structural response given partially observed data collected from a sensor network; the imputation process aims to firstly learn a factorized spatial feature U and a temporal feature X based on the observed data Y, and then reconstruct the response with imputed values; afterwards, given y:,t signifying the multivariate data at time t, the course of response forecasting utilizes the well-trained spatial factor U and the updated temporal factor X* to map L (≥1) historical sensing data to future T (≥1) structural responses, given by equation (1) which essentially establishes a temporal forecasting process; Section 3 of Page 11: the numerical analyses are performed on a standard PC with 28 Intel Core i9-7940X CPUs and 2 NVIDIA GTX 1080 Ti GPU; Section 3.1 of Page 12 with FIGS. 5(a)-(d) in Page 14: Figure 5(d) shows that each monitoring section S i i ∈ 1,2 has five strain sensors installed on the bottom of the hollow slab beam; vibrational chord strain gauges are installed which facilitate monitoring of both strain response of the bridge and the corresponding operation temperature (see Figure 5(b))) Kim and Ren are analogous art because they are from the same field of endeavor, a system and a method relating to Data Imputation and Response Forecasting. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of Ren to Kim. Motivation for doing so would expand the capability to various practical applications. Claims 2 and 10 Kim in view of Ren discloses all the elements as stated in Claims 1 and 9 respectively and further discloses wherein the first machine learning model comprises a fully connected neural network (Kim, Section IV.A-IV.B with FIG. 2 of Pages 8850-8852: present the architecture of imputation module that aims to impute missing values within the data; take advantage of GAN, so that our generator (G) learns the distribution of the real data under the supervision of the discriminator (D); particularly, we used Wasserstein GAN, which is easier to train stably than original GAN, alleviating the problem of non-convergence and mode dropping phenomenon; additionally, both G and D are gated recurrent unit (GRU)-based RNNs, where G is a bidirectional RNN with decaying cell; Since G takes as an incomplete time series data, it should handle two kinds of input variables, the observed variables and the missing data; for the former, where m i t = 1 , G works as an auto-encoder, which learns the distribution of the given data while reconstructing the input values; on the other hand, G does not reconstruct the same input value but estimates a new value to fill in the missing input value; the concept of denoising auto-encoder is to reconstruct a complete input data from a partially destroyed one, essentially learning to generate a clean “noise-reduced input”; based on this idea, we propose the dropping function, which helps the model to predict appropriate values for missing values by generating missing values that we know the ground truth: namely, X ~ ; given X ~ , not X, it is possible for the model to learn how to reconstruct the values of the unknown variables, because now there exists ground truth imputation labels for the missing values; dropping additional observed variables guarantees that missing values without ground truth are imputed appropriately and makes G learn how to remove noise in both observed variables and missing variables; to sum up, G behaves as an auto-encoder for the observed variable and authentic generator for the other; in order to indicate where missing values are, we concatenate time-series data X ~ and the mask matrix M ~ and then feed it to the generator; inspired by masked language modeling of BERT, we also make M ~ work as a mask token when G handles the missing values; with this in mind, we invert M ~   and then feed it to G, so that concatenated masking has the value of one at the observed and zero at the missing; additionally, missing data need to be imputed temporarily; we replace missing values with the mean value of each variable; finally, the input of G is shown as Z = X ~ ; ( 1 - M ~ ) ,   where ; denotes a concatenation operator; i.e., G receives additional information to distinguish observed variables from missing data through M ~ ; L2 loss is minimized to decrease the distance between prediction output of G and the raw timeseries data; finally, the loss of G for reconstructing input data is as shown in equation (3); In addition to loss LG, r, the generator is also trained by minimizing the adversarial loss, fooling D using the imputation results; the input of D is the observations with missing data replaced by the output of G; D learns to distinguish the observed variables from the estimated values, classifying the former as real and the latter as fake; specifically, input variables are classified elementwise; that is, the output of D is a d × T matrix whose elements are [Symbol font/0xCE] [0; 1]; we designed D to distinguish input sequence elementwise, as there rarely exists a sequence in which all variables are observed in highly-missing time-series data; with well-trained D, the G learns to produce complete data that follows the distribution of real data fooling D; altogether, the adversarial loss for G and D is shown as equations (4)-(6); the prevailing method for adopting RNN-based model to handle incomplete data is decaying input variable or hidden state vector of RNN cells using the time gap matrix δ i t ; if a variable is consecutively missing for a long time, the reliability of hidden state from the previous time step becomes low; therefore, the missing pattern of a variable with respect to time should be considered and the hidden state vector of RNN should be decayed if a variable has been missing for a long while; therefore, we propose a decaying method that feeds k time gap vectors, δ t - k + 1 : t , into a fully connected neural network with two linear layers followed by nonlinear functions shown in equations (7)-(8), where W η and W γ are model parameters to learn, k is a hyperparameter that denotes the length of time gap sequence and Maxpool indicates max pooling operation; Eq. (7) and Eq. (8) include a linear transformation that regards time gap variations between different variables and different time steps, respectively; with the assumption that there are multiple patterns of temporal time gap variations, we map η t to a two-dimensional space instead of a vector space and perform max pooling operation, i.e., selecting one pattern out of various candidates; accordingly, γ t is in the same vector space with the hidden state vector of RNN; introduce input decaying that downscales the input data x ~ t with the rate of ( 1 - γ t ) to keep the scale of gates in GRU constant; as hidden state vectors are directly related to predicting the imputation result, they should contain the information of underlying true distribution of input data and maintain a constant scale for its value; however, if the decaying rate is repeatedly multiplied, the hidden state vectors keep decreasing and the scale of the gates fluctuates; this inconsistency of the scale can lead to inappropriate prediction results; for this reason, we decay x ~ t with the rate of ( 1 - γ t )   so that the sum of coefficients of x ~ t and h t always equals to one; as a result, the update functions of GRU with the decaying mechanism is as follows: h t - 1 = γ t   ⨀   h t - 1 ,   z t = [ 1 - γ t ; 0 ] ⊙ z t ; u t = σ W u h t - 1 ; z t + b u ,   r t = σ W r h t - 1 ; z t + b r ; h ~ t = t a n h W h r t ⊙ h t - 1 ; z t + b h ; h t = 1 - u t ⊙ h t - 1 + u t ⊙ h ~ t , where W u , W r and W h are model parameters to learn; in Eq. (IV-B), we pad ( 1 - γ t ) with zero vectors to have the shape of z t ; note that it results in multiplying zero matrix and the concatenated m ~ t ). Claims 3 and 11 Kim in view of Ren discloses all the elements as stated in Claims 1 and 9 respectively and further discloses wherein the predicted future values are for a next step ahead and for multiple steps ahead (Kim, Section III.C in Page 8852-8853 with FIG. 2 in Page 8851 and FIG. 3 in Page 8853: introduce a novel deep learning model using the gating mechanism, gated classifier C, which considers the reliability of imputation results and substitutes the results partially with other values based on the reliability; it is proposed based on the question of suitability of imputation results based on auto-encoder or generative adversarial training for predicting downstream task; if the missing rate of input data is very high, there is little information that we can extract to estimate its distribution; if so, the output of the imputation module might not be reliable; since our prediction label is highly related to the true distribution, unreliable noisy outputs should not be used for downstream task; therefore, we propose gated classifier that mixes observations and predictions of the imputation module based on the reliability of each value and uses the resulting data for classification; gated classifier includes gating module which estimates the suitability of the output of the generator compared to the raw input data for predicting downstream task; Fig. 3 shows how a gating matrix is introduced; first, we define a gating matrix Λ ∈ R d × T that denotes the degree of reliability for Y, the output of the generator; the gating value, λ t i ∈ 0 ,   1 is an element of Λ   that quantifies the reliability of corresponding variable y t i ; we concatenate X ~ and Y to estimate relative confidence of Y compared to X ~ ; since there exists missing values substituted with the mean value in Y, we also concatenate inverted M ~ to indicate the location of missing values similarly to the formulation of generator’s input; note that the concatenation result X ~ T ; Y T ; 1 - M ~ T T is a 3d×T matrix; next, we compute the convolution of the concatenation with various kernel sizes; regardless of the kernel size, each convolution is a mapping R 3 d × T → R d × T , zero padding the input if necessary; i.e., the number of its output channels always equals the number of variables of X ~ ; each kernel sees the entire variable of raw data, generator output and inverted mask matrix at time t at once, so that gating module can make a decision about t-th variable based on the full information at time t; . As a consequence, the kernel size is (3d, *), where d decides how much temporal information is available; after convolutions, we conduct a max-pooling operation on convolution filter outputs so as to select one logit for each variable; as a result, our gating module output has the same dimension with raw data; finally, we apply sigmoid function to the output and it is used as a gating matrix; after we obtain the gating matrix, Λ , we mix the generator output and raw data by the ratio of gating value S = Y ⊙ Λ + X ~ ⊙ 1 - Λ , where S indicates the mixed output; it is what the gating module decides to be the best input data for performing downstream task; however, since we replaced missing values in X ~ with the mean value, which is zero due to the normalization, the mixing operation causes a difference in the scale of observations and the missing values where m ~ t i = 1 ; therefore, we compensate the shortfall with GRU cell weights; it should be filled by a ratio of (1 - λ t i ) for each missing values and no compensation is necessary for observed variables; thus, we multiply (1 - λ t i ) and (1 - M ~ ) to indicates the deficiency ratio to fill-in for each variable; and then, it is concatenated with S and fed to GRU; consequently, the first-layer weights of GRU are multiplied with the ratios and S with the former compensating the latter: W c s t ; λ t ∙ 1 - m ~ t ; the GRU layer of gated classifier is then followed by a fully-connected layer that outputs the classification results; the gated classifier and the imputation module, G and D, are jointly trained in an end-to-end manner; accordingly, with the assumption that the prediction label is strongly related to true distribution of the data, G generates better results that follow true distribution while minimizing classification loss) (Ren, Section 2 and Section 2.1 in Pages 3-4: formulate the problem of SHM data imputation and response forecasting in the context of incremental Bayesian tensor learning, and present the spatiotemporal dependency modeling procedure via matrix factorization; the goal of continuous/steaming SHM data imputation and forecasting is to estimate the missing values and predict the future structural response given partially observed data collected from a sensor network; the imputation process aims to firstly learn a factorized spatial feature U and a temporal feature X based on the observed data Y, and then reconstruct the response with imputed values; afterwards, given y:,t signifying the multivariate data at time t, the course of response forecasting utilizes the well-trained spatial factor U and the updated temporal factor X* to map L (≥1) historical sensing data to future T (≥1) structural responses, given by equation (1) which essentially establishes a temporal forecasting process;). Claims 4 and 12 Kim in view of Ren discloses all the elements as stated in Claims 1 and 9 respectively and further discloses wherein the second machine learning model comprises a convolutional neural network (Kim, Section III.C in Page 8852-8853 with FIG. 2 in Page 8851 and FIG. 3 in Page 8853: introduce a novel deep learning model using the gating mechanism, gated classifier C, which considers the reliability of imputation results and substitutes the results partially with other values based on the reliability; it is proposed based on the question of suitability of imputation results based on auto-encoder or generative adversarial training for predicting downstream task; if the missing rate of input data is very high, there is little information that we can extract to estimate its distribution; if so, the output of the imputation module might not be reliable; since our prediction label is highly related to the true distribution, unreliable noisy outputs should not be used for downstream task; therefore, we propose gated classifier that mixes observations and predictions of the imputation module based on the reliability of each value and uses the resulting data for classification; gated classifier includes gating module which estimates the suitability of the output of the generator compared to the raw input data for predicting downstream task; Fig. 3 shows how a gating matrix is introduced; first, we define a gating matrix Λ ∈ R d × T that denotes the degree of reliability for Y, the output of the generator; the gating value, λ t i ∈ 0 ,   1 is an element of Λ   that quantifies the reliability of corresponding variable y t i ; we concatenate X ~ and Y to estimate relative confidence of Y compared to X ~ ; since there exists missing values substituted with the mean value in Y, we also concatenate inverted M ~ to indicate the location of missing values similarly to the formulation of generator’s input; note that the concatenation result X ~ T ; Y T ; 1 - M ~ T T is a 3d×T matrix; next, we compute the convolution of the concatenation with various kernel sizes; regardless of the kernel size, each convolution is a mapping R 3 d × T → R d × T , zero padding the input if necessary; i.e., the number of its output channels always equals the number of variables of X ~ ; each kernel sees the entire variable of raw data, generator output and inverted mask matrix at time t at once, so that gating module can make a decision about t-th variable based on the full information at time t; . As a consequence, the kernel size is (3d, *), where d decides how much temporal information is available; after convolutions, we conduct a max-pooling operation on convolution filter outputs so as to select one logit for each variable; as a result, our gating module output has the same dimension with raw data; finally, we apply sigmoid function to the output and it is used as a gating matrix; after we obtain the gating matrix, Λ , we mix the generator output and raw data by the ratio of gating value S = Y ⊙ Λ + X ~ ⊙ 1 - Λ , where S indicates the mixed output; it is what the gating module decides to be the best input data for performing downstream task; however, since we replaced missing values in X ~ with the mean value, which is zero due to the normalization, the mixing operation causes a difference in the scale of observations and the missing values where m ~ t i = 1 ; therefore, we compensate the shortfall with GRU cell weights; it should be filled by a ratio of (1 - λ t i ) for each missing values and no compensation is necessary for observed variables; thus, we multiply (1 - λ t i ) and (1 - M ~ ) to indicates the deficiency ratio to fill-in for each variable; and then, it is concatenated with S and fed to GRU; consequently, the first-layer weights of GRU are multiplied with the ratios and S with the former compensating the latter: W c s t ; λ t ∙ 1 - m ~ t ; the GRU layer of gated classifier is then followed by a fully-connected layer that outputs the classification results; the gated classifier and the imputation module, G and D, are jointly trained in an end-to-end manner; accordingly, with the assumption that the prediction label is strongly related to true distribution of the data, G generates better results that follow true distribution while minimizing classification loss). Claims 5 and 13 Kim in view of Ren discloses all the elements as stated in Claims 4 and 12 respectively and further discloses wherein the convolutional neural network comprises linear layers and performs a depth-wise separable convolution on the different features (Kim, Section III.C in Page 8852-8853 with FIG. 2 in Page 8851 and FIG. 3 in Page 8853: introduce a novel deep learning model using the gating mechanism, gated classifier C, which considers the reliability of imputation results and substitutes the results partially with other values based on the reliability; it is proposed based on the question of suitability of imputation results based on auto-encoder or generative adversarial training for predicting downstream task; if the missing rate of input data is very high, there is little information that we can extract to estimate its distribution; if so, the output of the imputation module might not be reliable; since our prediction label is highly related to the true distribution, unreliable noisy outputs should not be used for downstream task; therefore, we propose gated classifier that mixes observations and predictions of the imputation module based on the reliability of each value and uses the resulting data for classification; gated classifier includes gating module which estimates the suitability of the output of the generator compared to the raw input data for predicting downstream task; Fig. 3 shows how a gating matrix is introduced; first, we define a gating matrix Λ ∈ R d × T that denotes the degree of reliability for Y, the output of the generator; the gating value, λ t i ∈ 0 ,   1 is an element of Λ   that quantifies the reliability of corresponding variable y t i ; we concatenate X ~ and Y to estimate relative confidence of Y compared to X ~ ; since there exists missing values substituted with the mean value in Y, we also concatenate inverted M ~ to indicate the location of missing values similarly to the formulation of generator’s input; note that the concatenation result X ~ T ; Y T ; 1 - M ~ T T is a 3d×T matrix; next, we compute the convolution of the concatenation with various kernel sizes; regardless of the kernel size, each convolution is a mapping R 3 d × T → R d × T , zero padding the input if necessary; i.e., the number of its output channels always equals the number of variables of X ~ ; each kernel sees the entire variable of raw data, generator output and inverted mask matrix at time t at once, so that gating module can make a decision about t-th variable based on the full information at time t; . As a consequence, the kernel size is (3d, *), where d decides how much temporal information is available; after convolutions, we conduct a max-pooling operation on convolution filter outputs so as to select one logit for each variable; as a result, our gating module output has the same dimension with raw data; finally, we apply sigmoid function to the output and it is used as a gating matrix; after we obtain the gating matrix, Λ , we mix the generator output and raw data by the ratio of gating value S = Y ⊙ Λ + X ~ ⊙ 1 - Λ , where S indicates the mixed output; it is what the gating module decides to be the best input data for performing downstream task; however, since we replaced missing values in X ~ with the mean value, which is zero due to the normalization, the mixing operation causes a difference in the scale of observations and the missing values where m ~ t i = 1 ; therefore, we compensate the shortfall with GRU cell weights; it should be filled by a ratio of (1 - λ t i ) for each missing values and no compensation is necessary for observed variables; thus, we multiply (1 - λ t i ) and (1 - M ~ ) to indicates the deficiency ratio to fill-in for each variable; and then, it is concatenated with S and fed to GRU; consequently, the first-layer weights of GRU are multiplied with the ratios and S with the former compensating the latter: W c s t ; λ t ∙ 1 - m ~ t ; the GRU layer of gated classifier is then followed by a fully-connected layer that outputs the classification results; the gated classifier and the imputation module, G and D, are jointly trained in an end-to-end manner; accordingly, with the assumption that the prediction label is strongly related to true distribution of the data, G generates better results that follow true distribution while minimizing classification loss). Claims 6 and 14 Kim in view of Ren discloses all the elements as stated in Claims 4 and 12 respectively and further discloses training the convolutional neural network to simultaneously learn the different functions for the corresponding different features (Kim, Section III with FIG. 1 of Page 8850: Let X n n = 1 N and l n n = 1 N denote a multivariate timeseries dataset and the corresponding label, respectively; X = x 1 ,   x 2 , … , x n ∈ X n n = 1 N with label l ∈ l n n = 1 N is a multivariate sequence of length T, where x t ∈ R d is the t-th observation of X at timestamp s t ; denote the i-th variable of the t-th observation as x i t and it is either observed or missing data; M = m 1 ,   m 2 , … , m T is a mask matrix whose element m i t ∈ 0,1 indicates whether the given variable x i t is missing or not; m i t = 1 if x i t is observed, otherwise m i t = 0 ; moreover, introduce a dropped matrix X ~ , which is generated by additionally dropping a portion of the observed variables from X to make model learn how to impute missing values; The corresponding mask matrix for X ~ , is denoted as M ~ ; define a time gap matrix Δ = δ 1 ,   δ 2 , … , δ T that serves time interval from the last observation for each variable of X ~ , i.e., δ i t shows how long a variable x ~ i has been missing consecutively until the t-th time step; Fig. 1 shows an example of model inputs with colors indicating the status of each variable: white for observed variables, red for missing variables and blue for additionally dropped variables, where X is a time series data of length five with three features; s is a time stamp vector for X; l is the given prediction label for X, which can be either classification or regression label; X ~ shows input data after additional 30% of the observed variables is dropped; M and Δ is a mask matrix and a time gap matrix for X ~ , respectively; Section IV.A-IV.B with FIG. 2 of Pages 8850-8852: present the architecture of imputation module that aims to impute missing values within the data; take advantage of GAN, so that our generator (G) learns the distribution of the real data under the supervision of the discriminator (D); particularly, we used Wasserstein GAN, which is easier to train stably than original GAN, alleviating the problem of non-convergence and mode dropping phenomenon; additionally, both G and D are gated recurrent unit (GRU)-based RNNs, where G is a bidirectional RNN with decaying cell; Since G takes as an incomplete time series data, it should handle two kinds of input variables, the observed variables and the missing data; for the former, where m i t = 1 , G works as an auto-encoder, which learns the distribution of the given data while reconstructing the input values; on the other hand, G does not reconstruct the same input value but estimates a new value to fill in the missing input value; the concept of denoising auto-encoder is to reconstruct a complete input data from a partially destroyed one, essentially learning to generate a clean “noise-reduced input”; based on this idea, we propose the dropping function, which helps the model to predict appropriate values for missing values by generating missing values that we know the ground truth: namely, X ~ ; given X ~ , not X, it is possible for the model to learn how to reconstruct the values of the unknown variables, because now there exists ground truth imputation labels for the missing values; dropping additional observed variables guarantees that missing values without ground truth are imputed appropriately and makes G learn how to remove noise in both observed variables and missing variables; to sum up, G behaves as an auto-encoder for the observed variable and authentic generator for the other; in order to indicate where missing values are, we concatenate time-series data X ~ and the mask matrix M ~ and then feed it to the generator; inspired by masked language modeling of BERT, we also make M ~ work as a mask token when G handles the missing values; with this in mind, we invert M ~   and then feed it to G, so that concatenated masking has the value of one at the observed and zero at the missing; additionally, missing data need to be imputed temporarily; we replace missing values with the mean value of each variable; finally, the input of G is shown as Z = X ~ ; ( 1 - M ~ ) ,   where ; denotes a concatenation operator; i.e., G receives additional information to distinguish observed variables from missing data through M ~ ; L2 loss is minimized to decrease the distance between prediction output of G and the raw timeseries data; finally, the loss of G for reconstructing input data is as shown in equation (3); In addition to loss LG, r, the generator is also trained by minimizing the adversarial loss, fooling D using the imputation results; the input of D is the observations with missing data replaced by the output of G; D learns to distinguish the observed variables from the estimated values, classifying the former as real and the latter as fake; specifically, input variables are classified elementwise; that is, the output of D is a d × T matrix whose elements are [Symbol font/0xCE] [0; 1]; we designed D to distinguish input sequence elementwise, as there rarely exists a sequence in which all variables are observed in highly-missing time-series data; with well-trained D, the G learns to produce complete data that follows the distribution of real data fooling D; altogether, the adversarial loss for G and D is shown as equations (4)-(6); the prevailing method for adopting RNN-based model to handle incomplete data is decaying input variable or hidden state vector of RNN cells using the time gap matrix δ i t ; if a variable is consecutively missing for a long time, the reliability of hidden state from the previous time step becomes low; therefore, the missing pattern of a variable with respect to time should be considered and the hidden state vector of RNN should be decayed if a variable has been missing for a long while; therefore, we propose a decaying method that feeds k time gap vectors, δ t - k + 1 : t , into a fully connected neural network with two linear layers followed by nonlinear functions shown in equations (7)-(8), where W η and W γ are model parameters to learn, k is a hyperparameter that denotes the length of time gap sequence and Maxpool indicates max pooling operation; Eq. (7) and Eq. (8) include a linear transformation that regards time gap variations between different variables and different time steps, respectively; with the assumption that there are multiple patterns of temporal time gap variations, we map η t to a two-dimensional space instead of a vector space and perform max pooling operation, i.e., selecting one pattern out of various candidates; accordingly, γ t is in the same vector space with the hidden state vector of RNN; introduce input decaying that downscales the input data x ~ t with the rate of ( 1 - γ t ) to keep the scale of gates in GRU constant; as hidden state vectors are directly related to predicting the imputation result, they should contain the information of underlying true distribution of input data and maintain a constant scale for its value; however, if the decaying rate is repeatedly multiplied, the hidden state vectors keep decreasing and the scale of the gates fluctuates; this inconsistency of the scale can lead to inappropriate prediction results; for this reason, we decay x ~ t with the rate of ( 1 - γ t )   so that the sum of coefficients of x ~ t and h t always equals to one; as a result, the update functions of GRU with the decaying mechanism is as follows: h t - 1 = γ t   ⨀   h t - 1 ,   z t = [ 1 - γ t ; 0 ] ⊙ z t ; u t = σ W u h t - 1 ; z t + b u ,   r t = σ W r h t - 1 ; z t + b r ; h ~ t = t a n h W h r t ⊙ h t - 1 ; z t + b h ; h t = 1 - u t ⊙ h t - 1 + u t ⊙ h ~ t , where W u , W r and W h are model parameters to learn; in Eq. (IV-B), we pad ( 1 - γ t ) with zero vectors to have the shape of z t ; note that it results in multiplying zero matrix and the concatenated m ~ t ; Section III.C in Page 8852-8853 with FIG. 2 in Page 8851 and FIG. 3 in Page 8853: introduce a novel deep learning model using the gating mechanism, gated classifier C, which considers the reliability of imputation results and substitutes the results partially with other values based on the reliability; it is proposed based on the question of suitability of imputation results based on auto-encoder or generative adversarial training for predicting downstream task; if the missing rate of input data is very high, there is little information that we can extract to estimate its distribution; if so, the output of the imputation module might not be reliable; since our prediction label is highly related to the true distribution, unreliable noisy outputs should not be used for downstream task; therefore, we propose gated classifier that mixes observations and predictions of the imputation module based on the reliability of each value and uses the resulting data for classification; gated classifier includes gating module which estimates the suitability of the output of the generator compared to the raw input data for predicting downstream task; Fig. 3 shows how a gating matrix is introduced; first, we define a gating matrix Λ ∈ R d × T that denotes the degree of reliability for Y, the output of the generator; the gating value, λ t i ∈ 0 ,   1 is an element of Λ   that quantifies the reliability of corresponding variable y t i ; we concatenate X ~ and Y to estimate relative confidence of Y compared to X ~ ; since there exists missing values substituted with the mean value in Y, we also concatenate inverted M ~ to indicate the location of missing values similarly to the formulation of generator’s input; note that the concatenation result X ~ T ; Y T ; 1 - M ~ T T is a 3d×T matrix; next, we compute the convolution of the concatenation with various kernel sizes; regardless of the kernel size, each convolution is a mapping R 3 d × T → R d × T , zero padding the input if necessary; i.e., the number of its output channels always equals the number of variables of X ~ ; each kernel sees the entire variable of raw data, generator output and inverted mask matrix at time t at once, so that gating module can make a decision about t-th variable based on the full information at time t; . As a consequence, the kernel size is (3d, *), where d decides how much temporal information is available; after convolutions, we conduct a max-pooling operation on convolution filter outputs so as to select one logit for each variable; as a result, our gating module output has the same dimension with raw data; finally, we apply sigmoid function to the output and it is used as a gating matrix; after we obtain the gating matrix, Λ , we mix the generator output and raw data by the ratio of gating value S = Y ⊙ Λ + X ~ ⊙ 1 - Λ , where S indicates the mixed output; it is what the gating module decides to be the best input data for performing downstream task; however, since we replaced missing values in X ~ with the mean value, which is zero due to the normalization, the mixing operation causes a difference in the scale of observations and the missing values where m ~ t i = 1 ; therefore, we compensate the shortfall with GRU cell weights; it should be filled by a ratio of (1 - λ t i ) for each missing values and no compensation is necessary for observed variables; thus, we multiply (1 - λ t i ) and (1 - M ~ ) to indicates the deficiency ratio to fill-in for each variable; and then, it is concatenated with S and fed to GRU; consequently, the first-layer weights of GRU are multiplied with the ratios and S with the former compensating the latter: W c s t ; λ t ∙ 1 - m ~ t ; the GRU layer of gated classifier is then followed by a fully-connected layer that outputs the classification results; the gated classifier and the imputation module, G and D, are jointly trained in an end-to-end manner; accordingly, with the assumption that the prediction label is strongly related to true distribution of the data, G generates better results that follow true distribution while minimizing classification loss). Claims 7 and 15 Kim in view of Ren discloses all the elements as stated in Claims 1 and 9 respectively and further discloses training the first and the second machine learning models simultaneously based on a combined loss (Kim, Section I of Page 8849: by jointly training the downstream task module and gating mechanism with adversarial loss, our model produces realistic and helpful imputation to predict the downstream task; a novel end-to-end model that imputes missing values and performs downstream tasks simultaneously with gating module and input dropping; the achievement of state-of-the- art performance on a real-world dataset; Section II of Page 8850: Cao et al. (2018) improved the downstream task prediction performance by training the imputation and the downstream task simultaneously, which greatly helped the model to generate appropriate imputation results which can also make high quality predictions for the downstream task; adopted GAN architecture that jointly train a discriminator and a generator on the multivariate time-series imputation task; Section III of Page 8850: addresses the problem of solving downstream tasks and imputing unknown variables within the given data jointly; in other words, predicting l and X using X ~ , M ~ and Δ ; propose an end-to-end GANs-based model that performs missing value imputation and downstream classification jointly; Section IV.A-IV.B with FIG. 2 of Pages 8850-8852: present the architecture of imputation module that aims to impute missing values within the data; take advantage of GAN, so that our generator (G) learns the distribution of the real data under the supervision of the discriminator (D); particularly, we used Wasserstein GAN, which is easier to train stably than original GAN, alleviating the problem of non-convergence and mode dropping phenomenon; additionally, both G and D are gated recurrent unit (GRU)-based RNNs, where G is a bidirectional RNN with decaying cell; Since G takes as an incomplete time series data, it should handle two kinds of input variables, the observed variables and the missing data; for the former, where m i t = 1 , G works as an auto-encoder, which learns the distribution of the given data while reconstructing the input values; on the other hand, G does not reconstruct the same input value but estimates a new value to fill in the missing input value; the concept of denoising auto-encoder is to reconstruct a complete input data from a partially destroyed one, essentially learning to generate a clean “noise-reduced input”; based on this idea, we propose the dropping function, which helps the model to predict appropriate values for missing values by generating missing values that we know the ground truth: namely, X ~ ; given X ~ , not X, it is possible for the model to learn how to reconstruct the values of the unknown variables, because now there exists ground truth imputation labels for the missing values; dropping additional observed variables guarantees that missing values without ground truth are imputed appropriately and makes G learn how to remove noise in both observed variables and missing variables; to sum up, G behaves as an auto-encoder for the observed variable and authentic generator for the other; in order to indicate where missing values are, we concatenate time-series data X ~ and the mask matrix M ~ and then feed it to the generator; inspired by masked language modeling of BERT, we also make M ~ work as a mask token when G handles the missing values; with this in mind, we invert M ~   and then feed it to G, so that concatenated masking has the value of one at the observed and zero at the missing; additionally, missing data need to be imputed temporarily; we replace missing values with the mean value of each variable; finally, the input of G is shown as Z = X ~ ; ( 1 - M ~ ) ,   where ; denotes a concatenation operator; i.e., G receives additional information to distinguish observed variables from missing data through M ~ ; L2 loss is minimized to decrease the distance between prediction output of G and the raw timeseries data; finally, the loss of G for reconstructing input data is as shown in equation (3); In addition to loss LG, r, the generator is also trained by minimizing the adversarial loss, fooling D using the imputation results; the input of D is the observations with missing data replaced by the output of G; D learns to distinguish the observed variables from the estimated values, classifying the former as real and the latter as fake; specifically, input variables are classified elementwise; that is, the output of D is a d × T matrix whose elements are [Symbol font/0xCE] [0; 1]; we designed D to distinguish input sequence elementwise, as there rarely exists a sequence in which all variables are observed in highly-missing time-series data; with well-trained D, the G learns to produce complete data that follows the distribution of real data fooling D; altogether, the adversarial loss for G and D is shown as equations (4)-(6); the prevailing method for adopting RNN-based model to handle incomplete data is decaying input variable or hidden state vector of RNN cells using the time gap matrix δ i t ; if a variable is consecutively missing for a long time, the reliability of hidden state from the previous time step becomes low; therefore, the missing pattern of a variable with respect to time should be considered and the hidden state vector of RNN should be decayed if a variable has been missing for a long while; therefore, we propose a decaying method that feeds k time gap vectors, δ t - k + 1 : t , into a fully connected neural network with two linear layers followed by nonlinear functions shown in equations (7)-(8), where W η and W γ are model parameters to learn, k is a hyperparameter that denotes the length of time gap sequence and Maxpool indicates max pooling operation; Eq. (7) and Eq. (8) include a linear transformation that regards time gap variations between different variables and different time steps, respectively; with the assumption that there are multiple patterns of temporal time gap variations, we map η t to a two-dimensional space instead of a vector space and perform max pooling operation, i.e., selecting one pattern out of various candidates; accordingly, γ t is in the same vector space with the hidden state vector of RNN; introduce input decaying that downscales the input data x ~ t with the rate of ( 1 - γ t ) to keep the scale of gates in GRU constant; as hidden state vectors are directly related to predicting the imputation result, they should contain the information of underlying true distribution of input data and maintain a constant scale for its value; however, if the decaying rate is repeatedly multiplied, the hidden state vectors keep decreasing and the scale of the gates fluctuates; this inconsistency of the scale can lead to inappropriate prediction results; for this reason, we decay x ~ t with the rate of ( 1 - γ t )   so that the sum of coefficients of x ~ t and h t always equals to one; as a result, the update functions of GRU with the decaying mechanism is as follows: h t - 1 = γ t   ⨀   h t - 1 ,   z t = [ 1 - γ t ; 0 ] ⊙ z t ; u t = σ W u h t - 1 ; z t + b u ,   r t = σ W r h t - 1 ; z t + b r ; h ~ t = t a n h W h r t ⊙ h t - 1 ; z t + b h ; h t = 1 - u t ⊙ h t - 1 + u t ⊙ h ~ t , where W u , W r and W h are model parameters to learn; in Eq. (IV-B), we pad ( 1 - γ t ) with zero vectors to have the shape of z t ; note that it results in multiplying zero matrix and the concatenated m ~ t ; Section III.C in Page 8852-8853 with FIG. 2 in Page 8851 and FIG. 3 in Page 8853: introduce a novel deep learning model using the gating mechanism, gated classifier C, which considers the reliability of imputation results and substitutes the results partially with other values based on the reliability; it is proposed based on the question of suitability of imputation results based on auto-encoder or generative adversarial training for predicting downstream task; if the missing rate of input data is very high, there is little information that we can extract to estimate its distribution; if so, the output of the imputation module might not be reliable; since our prediction label is highly related to the true distribution, unreliable noisy outputs should not be used for downstream task; therefore, we propose gated classifier that mixes observations and predictions of the imputation module based on the reliability of each value and uses the resulting data for classification; gated classifier includes gating module which estimates the suitability of the output of the generator compared to the raw input data for predicting downstream task; Fig. 3 shows how a gating matrix is introduced; first, we define a gating matrix Λ ∈ R d × T that denotes the degree of reliability for Y, the output of the generator; the gating value, λ t i ∈ 0 ,   1 is an element of Λ   that quantifies the reliability of corresponding variable y t i ; we concatenate X ~ and Y to estimate relative confidence of Y compared to X ~ ; since there exists missing values substituted with the mean value in Y, we also concatenate inverted M ~ to indicate the location of missing values similarly to the formulation of generator’s input; note that the concatenation result X ~ T ; Y T ; 1 - M ~ T T is a 3d×T matrix; next, we compute the convolution of the concatenation with various kernel sizes; regardless of the kernel size, each convolution is a mapping R 3 d × T → R d × T , zero padding the input if necessary; i.e., the number of its output channels always equals the number of variables of X ~ ; each kernel sees the entire variable of raw data, generator output and inverted mask matrix at time t at once, so that gating module can make a decision about t-th variable based on the full information at time t; . As a consequence, the kernel size is (3d, *), where d decides how much temporal information is available; after convolutions, we conduct a max-pooling operation on convolution filter outputs so as to select one logit for each variable; as a result, our gating module output has the same dimension with raw data; finally, we apply sigmoid function to the output and it is used as a gating matrix; after we obtain the gating matrix, Λ , we mix the generator output and raw data by the ratio of gating value S = Y ⊙ Λ + X ~ ⊙ 1 - Λ , where S indicates the mixed output; it is what the gating module decides to be the best input data for performing downstream task; however, since we replaced missing values in X ~ with the mean value, which is zero due to the normalization, the mixing operation causes a difference in the scale of observations and the missing values where m ~ t i = 1 ; therefore, we compensate the shortfall with GRU cell weights; it should be filled by a ratio of (1 - λ t i ) for each missing values and no compensation is necessary for observed variables; thus, we multiply (1 - λ t i ) and (1 - M ~ ) to indicates the deficiency ratio to fill-in for each variable; and then, it is concatenated with S and fed to GRU; consequently, the first-layer weights of GRU are multiplied with the ratios and S with the former compensating the latter: W c s t ; λ t ∙ 1 - m ~ t ; the GRU layer of gated classifier is then followed by a fully-connected layer that outputs the classification results; the gated classifier and the imputation module, G and D, are jointly trained in an end-to-end manner; accordingly, with the assumption that the prediction label is strongly related to true distribution of the data, G generates better results that follow true distribution while minimizing classification loss; Section IV of Page 8850 with FIG. 2 in Page 8851: model is designed for end-to-end missing value imputation and downstream task prediction in multivariate time-series data; the model consists of two modules: the imputation and the prediction module; Fig. 2 shows an overview of the proposed model, wherein y t and p t denotes the prediction output of the generator and the discriminator at time t, respectively; l ^ indicates the prediction output of the gated classifier for downstream task; the proposed decaying mechanism explained is adopted to the generator with its cells receiving δ t ; on the other hand, the discriminator and the gated classifier is GRU without any decaying). Claims 8 and 16 Kim in view of Ren discloses all the elements as stated in Claims 1 and 9 respectively and further discloses training the first and the second machine learning models via: determining an imputation loss for the first machine learning model imputing missing data values for training data (Kim, Section IV.A-IV.B with FIG. 2 of Pages 8850-8852: present the architecture of imputation module that aims to impute missing values within the data; take advantage of GAN, so that our generator (G) learns the distribution of the real data under the supervision of the discriminator (D); particularly, we used Wasserstein GAN, which is easier to train stably than original GAN, alleviating the problem of non-convergence and mode dropping phenomenon; additionally, both G and D are gated recurrent unit (GRU)-based RNNs, where G is a bidirectional RNN with decaying cell; Since G takes as an incomplete time series data, it should handle two kinds of input variables, the observed variables and the missing data; for the former, where m i t = 1 , G works as an auto-encoder, which learns the distribution of the given data while reconstructing the input values; on the other hand, G does not reconstruct the same input value but estimates a new value to fill in the missing input value; the concept of denoising auto-encoder is to reconstruct a complete input data from a partially destroyed one, essentially learning to generate a clean “noise-reduced input”; based on this idea, we propose the dropping function, which helps the model to predict appropriate values for missing values by generating missing values that we know the ground truth: namely, X ~ ; given X ~ , not X, it is possible for the model to learn how to reconstruct the values of the unknown variables, because now there exists ground truth imputation labels for the missing values; dropping additional observed variables guarantees that missing values without ground truth are imputed appropriately and makes G learn how to remove noise in both observed variables and missing variables; to sum up, G behaves as an auto-encoder for the observed variable and authentic generator for the other; in order to indicate where missing values are, we concatenate time-series data X ~ and the mask matrix M ~ and then feed it to the generator; inspired by masked language modeling of BERT, we also make M ~ work as a mask token when G handles the missing values; with this in mind, we invert M ~   and then feed it to G, so that concatenated masking has the value of one at the observed and zero at the missing; additionally, missing data need to be imputed temporarily; we replace missing values with the mean value of each variable; finally, the input of G is shown as Z = X ~ ; ( 1 - M ~ ) ,   where ; denotes a concatenation operator; i.e., G receives additional information to distinguish observed variables from missing data through M ~ ; L2 loss is minimized to decrease the distance between prediction output of G and the raw timeseries data; finally, the loss of G for reconstructing input data is as shown in equation (3); In addition to loss LG, r, the generator is also trained by minimizing the adversarial loss, fooling D using the imputation results; the input of D is the observations with missing data replaced by the output of G; D learns to distinguish the observed variables from the estimated values, classifying the former as real and the latter as fake; specifically, input variables are classified elementwise; that is, the output of D is a d × T matrix whose elements are [Symbol font/0xCE] [0; 1]; we designed D to distinguish input sequence elementwise, as there rarely exists a sequence in which all variables are observed in highly-missing time-series data; with well-trained D, the G learns to produce complete data that follows the distribution of real data fooling D; altogether, the adversarial loss for G and D is shown as equations (4)-(6); the prevailing method for adopting RNN-based model to handle incomplete data is decaying input variable or hidden state vector of RNN cells using the time gap matrix δ i t ; if a variable is consecutively missing for a long time, the reliability of hidden state from the previous time step becomes low; therefore, the missing pattern of a variable with respect to time should be considered and the hidden state vector of RNN should be decayed if a variable has been missing for a long while; therefore, we propose a decaying method that feeds k time gap vectors, δ t - k + 1 : t , into a fully connected neural network with two linear layers followed by nonlinear functions shown in equations (7)-(8), where W η and W γ are model parameters to learn, k is a hyperparameter that denotes the length of time gap sequence and Maxpool indicates max pooling operation; Eq. (7) and Eq. (8) include a linear transformation that regards time gap variations between different variables and different time steps, respectively; with the assumption that there are multiple patterns of temporal time gap variations, we map η t to a two-dimensional space instead of a vector space and perform max pooling operation, i.e., selecting one pattern out of various candidates; accordingly, γ t is in the same vector space with the hidden state vector of RNN; introduce input decaying that downscales the input data x ~ t with the rate of ( 1 - γ t ) to keep the scale of gates in GRU constant; as hidden state vectors are directly related to predicting the imputation result, they should contain the information of underlying true distribution of input data and maintain a constant scale for its value; however, if the decaying rate is repeatedly multiplied, the hidden state vectors keep decreasing and the scale of the gates fluctuates; this inconsistency of the scale can lead to inappropriate prediction results; for this reason, we decay x ~ t with the rate of ( 1 - γ t )   so that the sum of coefficients of x ~ t and h t always equals to one; as a result, the update functions of GRU with the decaying mechanism is as follows: h t - 1 = γ t   ⨀   h t - 1 ,   z t = [ 1 - γ t ; 0 ] ⊙ z t ; u t = σ W u h t - 1 ; z t + b u ,   r t = σ W r h t - 1 ; z t + b r ; h ~ t = t a n h W h r t ⊙ h t - 1 ; z t + b h ; h t = 1 - u t ⊙ h t - 1 + u t ⊙ h ~ t , where W u , W r and W h are model parameters to learn; in Eq. (IV-B), we pad ( 1 - γ t ) with zero vectors to have the shape of z t ; note that it results in multiplying zero matrix and the concatenated m ~ t ); determining a prediction loss for the second machine learning model predicting future data values related to the training data; and combining the imputation loss and the prediction loss to determine an overall loss, wherein the first and the second machine learning models are trained simultaneously based on the overall loss (Kim, Section I of Page 8849: by jointly training the downstream task module and gating mechanism with adversarial loss, our model produces realistic and helpful imputation to predict the downstream task; a novel end-to-end model that imputes missing values and performs downstream tasks simultaneously with gating module and input dropping; the achievement of state-of-the- art performance on a real-world dataset; Section II of Page 8850: Cao et al. (2018) improved the downstream task prediction performance by training the imputation and the downstream task simultaneously, which greatly helped the model to generate appropriate imputation results which can also make high quality predictions for the downstream task; adopted GAN architecture that jointly train a discriminator and a generator on the multivariate time-series imputation task; Section III of Page 8850: addresses the problem of solving downstream tasks and imputing unknown variables within the given data jointly; in other words, predicting l and X using X ~ , M ~ and Δ ; propose an end-to-end GANs-based model that performs missing value imputation and downstream classification jointly; Section III.C in Page 8852-8853 with FIG. 2 in Page 8851 and FIG. 3 in Page 8853: introduce a novel deep learning model using the gating mechanism, gated classifier C, which considers the reliability of imputation results and substitutes the results partially with other values based on the reliability; it is proposed based on the question of suitability of imputation results based on auto-encoder or generative adversarial training for predicting downstream task; if the missing rate of input data is very high, there is little information that we can extract to estimate its distribution; if so, the output of the imputation module might not be reliable; since our prediction label is highly related to the true distribution, unreliable noisy outputs should not be used for downstream task; therefore, we propose gated classifier that mixes observations and predictions of the imputation module based on the reliability of each value and uses the resulting data for classification; gated classifier includes gating module which estimates the suitability of the output of the generator compared to the raw input data for predicting downstream task; Fig. 3 shows how a gating matrix is introduced; first, we define a gating matrix Λ ∈ R d × T that denotes the degree of reliability for Y, the output of the generator; the gating value, λ t i ∈ 0 ,   1 is an element of Λ   that quantifies the reliability of corresponding variable y t i ; we concatenate X ~ and Y to estimate relative confidence of Y compared to X ~ ; since there exists missing values substituted with the mean value in Y, we also concatenate inverted M ~ to indicate the location of missing values similarly to the formulation of generator’s input; note that the concatenation result X ~ T ; Y T ; 1 - M ~ T T is a 3d×T matrix; next, we compute the convolution of the concatenation with various kernel sizes; regardless of the kernel size, each convolution is a mapping R 3 d × T → R d × T , zero padding the input if necessary; i.e., the number of its output channels always equals the number of variables of X ~ ; each kernel sees the entire variable of raw data, generator output and inverted mask matrix at time t at once, so that gating module can make a decision about t-th variable based on the full information at time t; . As a consequence, the kernel size is (3d, *), where d decides how much temporal information is available; after convolutions, we conduct a max-pooling operation on convolution filter outputs so as to select one logit for each variable; as a result, our gating module output has the same dimension with raw data; finally, we apply sigmoid function to the output and it is used as a gating matrix; after we obtain the gating matrix, Λ , we mix the generator output and raw data by the ratio of gating value S = Y ⊙ Λ + X ~ ⊙ 1 - Λ , where S indicates the mixed output; it is what the gating module decides to be the best input data for performing downstream task; however, since we replaced missing values in X ~ with the mean value, which is zero due to the normalization, the mixing operation causes a difference in the scale of observations and the missing values where m ~ t i = 1 ; therefore, we compensate the shortfall with GRU cell weights; it should be filled by a ratio of (1 - λ t i ) for each missing values and no compensation is necessary for observed variables; thus, we multiply (1 - λ t i ) and (1 - M ~ ) to indicates the deficiency ratio to fill-in for each variable; and then, it is concatenated with S and fed to GRU; consequently, the first-layer weights of GRU are multiplied with the ratios and S with the former compensating the latter: W c s t ; λ t ∙ 1 - m ~ t ; the GRU layer of gated classifier is then followed by a fully-connected layer that outputs the classification results; the gated classifier and the imputation module, G and D, are jointly trained in an end-to-end manner; accordingly, with the assumption that the prediction label is strongly related to true distribution of the data, G generates better results that follow true distribution while minimizing classification loss; Section IV of Page 8850 with FIG. 2 in Page 8851: model is designed for end-to-end missing value imputation and downstream task prediction in multivariate time-series data; the model consists of two modules: the imputation and the prediction module; Fig. 2 shows an overview of the proposed model, wherein y t and p t denotes the prediction output of the generator and the discriminator at time t, respectively; l ^ indicates the prediction output of the gated classifier for downstream task; the proposed decaying mechanism explained is adopted to the generator with its cells receiving δ t ; on the other hand, the discriminator and the gated classifier is GRU without any decaying). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Zuo et al. ("Graph Convolutional Networks for Traffic Forecasting with Missing Values", arXiv:2212.06419v1, Dec. 13, 2022, pp. 1-30) discloses in ABSTRACT of Page 1 that (1) traffic data usually contains missing values due to sensor or communication errors; (2) the Spatio-temporal feature in traffic data brings more challenges for processing such missing values, for which the classic techniques (e.g., data imputations) are limited: a) in temporal axis, the values can be randomly or consecutively missing; and b) in spatial axis, the missing values can happen on one single sensor or on multiple sensors simultaneously; (3) recent models powered by Graph Neural Networks achieved satisfying performance on traffic forecasting tasks; (4) however, few of them are applicable to such a complex missing-value context; (5) propose GCN-M, a Graph Convolutional Network model with the ability to handle the complex missing values in the Spatio-temporal context; (6) particularly, jointly model the missing value processing and traffic forecasting tasks, considering both local Spatio-temporal features and global historical patterns in an attention-based memory network; (7) propose as well a dynamic graph learning module based on the learned local-global features; and (8) the experimental results on real-life datasets show the reliability of our proposed method. Zuo further discloses in Section 1 with FIG. 1 of Pages 1-4 that (1) traffic forecasting has played a critical role in intelligent transportation systems, which helps the transportation department better manage and control traffic congestion; (2) generally represented by geo-located Multivariate Time Series (MTS), traffic data not only shows the typical characteristics of MTS, i.e., temporal dependency (Zuo et al, 2021), but also integrates the spatial information of the traffic network, i.e., the spatial dependency between the sensor traffic nodes over the road network; (3) since the traffic data is generally collected from geolocated sensors, sensor failures or communication errors will result in missing values in the collected data, thus deteriorating the performance of the forecasting model; (4) we should remark that the missing measures are usually marked as zero in traffic data (Li et al, 2021), which should be distinguished from the non-missing measures but with zero values; (5) the missing values can either be ignored in the learning model when calculating the loss function (Wang et al, 2020) or be considered before or during the training process (Cui et al, 2020b); (6) ignoring the missing values, especially when the missing ratio is high (Cui et al, 2020b), hinders the model from benefiting from the rich data information for better performance; (7) when considering the missing values in traffic data, most work (Cirstea et al, 2019) conducts data imputation during the preprocessing step, then imports the completed data into the training step, i.e., two-step processing; (8) recent work tends to jointly consider the missing values and the forecasting modeling during the training step (i.e., one-step processing) and declared better performance than the two-step processing (Che et al, 2018; Cui et al, 2020a,b; Tian et al, 2018; Tang et al, 2020); (9) however, the above-mentioned work suffers from three major issues: (a) first, the missing and zero values are usually considered to be the same, leading to unnecessary, even harmful data imputations, thus contradicting the raw data information; (b) second, most of the work (Che et al, 2018; Cui et al, 2020a; Tian et al, 2018; Tang et al, 2020) considers missing values from the temporal aspect, ignoring the rich information from the spatial perspective; and (c) third, they are generally designed for processing the missing values in some basic scenarios, such as random missing values or temporal block missing values, but lack power for the complex scenarios as shown in Fig. 1; (10) in the real world, the missing values in traffic data occur in both long-range (e.g., device power-off ) and short-range (e.g. device errors) settings, in partial (e.g., local sensor errors) and entire transportation network (e.g., control center errors); (11) therefore, a holistic approach is required for handling various types of missing values together in complex scenarios; (12) to handle both the Spatio-temporal patterns and complex missing-value scenarios in traffic data, propose Graph Convolutional Networks for Traffic Forecasting with Missing Values (GCN-M); (13) the graph neural network-based structure allows jointly modeling the Spatio-temporal patterns and the missing values in a one-step process; (14) construct local statistical features from spatial and temporal perspectives for handling short-range missing values; (15) this is further enhanced by a memory module to extract global historical features for processing long-range missing blocks; (16) the combined local-global features allow not only for identifying the missing measures from the inherent zero values but also for enriching the traffic embeddings, thus generating dynamic traffic graphs to model the dynamic spatial interactions between traffic nodes; (17) the missing values on a partial and entire network can then be considered from spatial and temporal perspectives; and (18) summarize the paper's main contributions as follows: (a) Complex missing value modeling: study the complex scenario where missing traffic values occur on both short & long ranges and on partial & entire transportation networks; (b) Spatio-temporal memory module: propose a memory module that can be used by GCN-M to learn both local Spatio-temporal features and global historical patterns in traffic data for handling the complex missing values; (c) Dynamic graph modeling: propose a dynamic graph convolution module that models the dynamic spatial interactions, wherein the dynamic graph is characterized by the learned local-global features at each timestamp, which not only offset the missing values' impact but also help learn the graph; (d) Joint model optimization: jointly model the Spatio-temporal patterns and missing values in one-step processing, which allows processing missing values specifically for traffic forecasting tasks, thus bringing better model performance than two-step processing; and (e) Extensive experiments on real-life data: the experiments are carried out on two real-life traffic datasets, wherein we provide detailed evaluations with 12 baselines, which show the effectiveness of GCN-M over state-of-the-art. Zuo also discloses in Section 3 in Pages 5-6 that (1) predict future traffic data by leveraging historical traffic data; (2) traffic data can be represented as a multivariate time series on a traffic network; (3) the traffic network includes a set of N traffic sensor nodes and a set of E edges connecting the nodes; (4) each node contains F features representing traffic flow, speed, occupancy, etc.; and (5) aim to build a model f, which can take an incomplete traffic sequence and the traffic network as input, to predict the traffic data for the next Tp time steps. Zuo further teaches in Section 4 with FIGS. 2-5 in Pages 6-13 that (1) traffic data is collected under complex urban conditions; (2) apart from the Spatio-temporal patterns in the traffic data, we also consider the scenarios of complex missing values; (3) we design a solution that models the local Spatiotemporal features and global historical patterns in a dynamic manner; (4) the complex missing values are considered when building the forecasting model, i.e., one-step processing; (5) the global structure of GCN-M is shown in Fig. 2, integrating a Multi-scale Memory Network module, an Output Forecasting module, and l Spatio-Temporal (ST) blocks. Each ST block integrates three key components: Temporal Convolution, Dynamic Graph Construction, and Dynamic Graph Convolution; (6) the input traffic observations and the mask sequence are fed into the multi-scale memory network to extract the local statistic features and global historical patterns thus enriching the traffic embeddings; (7) on the one hand, the enriched embeddings on each ST block are used to mark the dynamic traffic status, thus generating dynamic graphs by combining both static node embeddings and predefined graph information: (8) on the other hand, the learned dynamic graphs are combined with the temporal convolution module via a dynamic graph convolution to capture temporal and spatial dependencies in the traffic embeddings; (9) we adopt residual connections between the input and output of each ST block to avoid the gradient vanishing problem; (10) the output forecasting module takes the skip connections on the output of the final ST block and the hidden states after each temporal convolution for final prediction; (11) to extract the local statistic features and global historical patterns then form an enriched embedding, we adopt the concept of memory network, which was firstly proposed in (Weston et al, 2015) with primary application in Question-Answer (QA) systems; (12) as shown in Fig. 3, the main idea of our memory network is to learn from historical memory components which conserve the long-range multi-scale patterns, i.e., recent, daily-periodic, and weekly-periodic dependencies; (13) the scale range depends on the data characteristics; (14) specifically, we first extract local Spatio-temporal features as keys to query the memory components, the weighted historical long-range patterns will be cooperated with the local statistic features to eliminate the side effect from the missing values, and then, the local-global features will be output as the enriched traffic embeddings; (15) we first extract the Spatio-temporal features using the contextual information from observed parts of the time series; (16) consider both temporal and spatial aspects for generating the following statistic features of every timestamp: (a) Empirical Temporal Mean: the mean of previous observations reflects the recent traffic state and serves as a contextual knowledge; (b) Last Temporal Observation: adopt the assumption in (Che et al, 2018) that any missing value inherits more or less the information from the last non-missing observation; i.e., the temporal neighbor stays close to the current missing value; (c) Empirical Spatial Mean: another contextual knowledge is from the nearby nodes, which reflects the current local traffic situation; and (d) Nearest Spatial Observation: typically, the state of a graph node remains relatively similar to its neighbors, especially in a traffic graph where the nearby nodes share similar traffic situations; (17) global historical patterns play a critical role in building an enriched traffic embedding; (18) the historical observations in multiple scales (e.g., hourly, daily, weekly) can be embedded into memory as complement information for the local features; (18) the main idea is to adopt local features to query similar historical patterns in the memory and output a weighted feature representation for the current timestamp; (19) in this manner, the enriched multi-scale historical and local features allow not only eliminating the side effect of missing values but also improving the current feature embeddings; (19) in the embedding space, we compute the attention score between the query and each memory by taking the inner product followed by a Softmax; (20) the attention score represents the similarity of each historical observation to the query; (21) any pattern with a higher attention score is more similar to the context of targeting missing values; (22) as shown in Fig. 3, the response vector from memory is then a sum over the output memory vectors, weighted by the attention score from the input; (23) we can finally integrate both local Spatio-temporal and global multi-scale features and output the enriched traffic embeddings; (24) construct dynamic graphs (i.e., adjacency matrix) with the enriched traffic embeddings at each ST block, which integrates both local and global multi-scale patterns at each time step, which allows capturing the spatial relationship between traffic nodes robustly; and (25) as shown in Fig. 4, the main idea here is to generate dynamic filters from the predefined graphs and the traffic embeddings, which are applied on the randomly initialized static node embeddings to construct dynamic adjacency matrices; (26) the temporal convolution network (TCN) (Lea et al, 2017) consists of multiple dilated convolution layers, which allows extracting high-level temporal trends; (27) as shown in Fig. 5, considering the temporal dynamics in traffic data, we adopt the temporal convolution module (Wu et al, 2019) with the consideration of the gating mechanism over the enriched traffic embeddings; (28) one dilated convolution block is followed by a tangent hyperbolic activation function to output the temporal features; (29) the other block is followed by a sigmoid activation function as a gate to determine the ratio of information that can pass to the next module; (30) in particular, the sigmoid gate controls which input of the current states is relevant for discovering compositional structure and dynamic variances in time series; (31) a classic temporal convolution module stacks the temporal features at each time step t; (32) therefore, the upper layer contains richer information than the lower layer; (33) the gating mechanism allows filtering the temporal features on the lower layers by weighting features on different time steps without considering the spatial node interactions at each time step. Moreover, the spatial interactions in traffic data always show a dynamic nature (Wu et al, 2020); (34) to this end, the gating mechanism from a dynamic spatial aspect is envisaged to better capture the Spatio-temporal patterns; (35) spatial interactions between the traffic nodes could be used to improve traffic forecasting performance; (36) the dynamic spatial interaction leads to considering a dynamic version of graph convolution to conduct it on different graphs at different timestamps; (37) adopt the enriched traffic embeddings, which consider the missing-value issues to generate robust dynamic graphs; (38) the outputs of the middle temporal convolution modules and the enriched traffic embeddings of the last ST block are considered for the final prediction, which represent the hidden states at various Spatio-temporal levels; (39) two fully-connected layers are added to project the concatenated features into the desired output dimension; and (40) given the ground truth and the predictions, we use mean absolute error (MAE) as our model's loss function for training. Any inquiry concerning this communication or earlier communications from the examiner should be directed to HWEI-MIN LU whose telephone number is (313)446-4913. The examiner can normally be reached Mon - Fri: 9:00 AM - 6:00 PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela D. Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HWEI-MIN LU/Primary Examiner, Art Unit 2142
Read full office action

Prosecution Timeline

Jan 05, 2024
Application Filed
Jul 23, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705533
SYSTEMS AND METHODS FOR IMPROVING PREDICTION PROCESS USING AUTOMATED RULE LEARNING FRAMEWORK
3y 8m to grant Granted Aug 11, 2026
Patent 12700003
SYSTEMS AND METHODS FOR FREQUENT MACHINE LEARNING MODEL RETRAINING AND RULE OPTIMIZATION
4y 2m to grant Granted Aug 04, 2026
Patent 12694335
SYSTEMS AND METHODS FOR REPURPOSING A MACHINE LEARNING MODEL
3y 6m to grant Granted Jul 28, 2026
Patent 12682260
ADJUDICATION ALGORITHM BYPASS CONDITIONS
4y 1m to grant Granted Jul 14, 2026
Patent 12675690
ANOMALY DETECTION WITH MODEL HYPERPARAMETER SELECTION
4y 3m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
63%
Grant Probability
99%
With Interview (+39.6%)
2y 11m (~3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 233 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month