Prosecution Insights
Last updated: October 01, 2026
Application No. 17/677,839

METHODS AND APPARATUS FOR DATA IMPUTATION OF A SPARSE TIME SERIES DATA SET

Final Rejection §103
Filed
Feb 22, 2022
Priority
Aug 24, 2021 — IN 202141038261
Examiner
PHAKOUSONH, DARAVANH
Art Unit
2121
Tech Center
2100 — Computer Architecture & Software
Assignee
Walmart Apollo LLC
OA Round
4 (Final)
25%
Grant Probability
At Risk
5-6
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants only 25% of cases
25%
Career Allowance Rate
1 granted / 4 resolved
-30.0% vs TC avg
Strong +100% interview lift
Without
With
+100.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
24 currently pending
Career history
41
Total Applications
across all art units

Statute-Specific Performance

§101
52.8%
+12.8% vs TC avg
§103
13.7%
-26.3% vs TC avg
§102
19.9%
-20.1% vs TC avg
§112
12.4%
-27.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 4 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment/Arguments 1. Applicant’s amendments to the claims overcomes the rejection under 35 U.S.C. 103. 2. Applicant’s arguments filed on June 25, 2026, regarding the rejection under 35 U.S.C. 103 have been fully considered but are not persuasive. Applicant’s arguments regarding the rejection of independent claims 1, 11, and 20 have been fully considered but are not persuasive. Applicant first identifies the limitations directed to executing a recurrent neural network using the first, second, and third data sets to iteratively generate substitute retail order data. These limitations have been reconsidered and remain taught by the combined references for the reasons set forth in the rejection of claim 1. Claim 11 and 20 recite corresponding limitations and are rejected for the same reasons identified in their respective rejections. Applicant argues that Fouladgar cannot be combined with Chen because Fouladgar allegedly teaches that missing data patterns from one domain cannot be applied to another domain. The cited portions of Fouladgar do not contain that teaching. Fouladgar’s conclusion was only that “not all missing patterns provide meaningful representation of missingness” based on its evaluation of different missingness information in the meteorological data sets. Fouladgar did not conclude that an imputation framework developed or evaluated in one domain cannot be applied to another domain. To the contrary, Fouladgar proposes replicating the experiments using a health-related daily stress monitoring data set. Although this statement does not establish the results of the proposed testing, it demonstrates that Fouladgar contemplated evaluating the same approach using time series data from another domain. Accordingly, Fouladgar does not warn against applying its recurrent imputation framework to other time series data or indicate that applying the framework to Chen’s retail time series data would be unsuccessful. Applicant correctly notes that Fouladgar’s reported experiments involved meteorological multivariate time series data, including Beijing PM2.5, Italy Air Quality, and Beijing Mult-Site Air Quality data sets. The rejection, however, does not rely on Fouladgar alone for the claimed aggregate retail order data. As explained in the rejection of claim 1, Fouladgar provides a recurrent neural network framework for imputing missing values in multivariate time series data, while Chen provides retail time series data, fitting model parameters using historical retail data, and generating retail forecasts. Obviousness does not require each reference to individually disclose every limitation of the claim. The proposed combination does not require substituting meteorological observations for retail order data. Rather, it applies Fouladgar’s recurrent imputation framework to the retail time series data taught by Chen. Meteorological data and retail data may both be represented as multivariate time series data containing historical observations, seasonal patterns, multiple related variables, missing values, and unusual events. A person of ordinary skill in the art therefore would have recognized that Fouladgar’s technique for reconstructing incomplete multivariate time series data could be applied to Chen’s incomplete retail time series data before fitting a model and generating a retail forecast. Cao further supports applying time series imputation methods across different domains. Cao states that multivariate time series data are used in financial marketing, health care, meteorology, and traffic engineering and that missing values may significantly harm downstream classification and regression applications. Cao therefore identifies missing value imputation as a problem shared by multivariate time series data from different fields, rather than a problem restricted to meteorological data. This further supports applying the imputation methods of Fouladgar and Cao to the retail time series data taught by Chen. The complete reason for combining the references is provided in the rejection of claim 1. The rejection of claim 1 also addresses the limitations added by amendment. As set forth in the rejection, Cao provides the forward and backward recurrent computations, passage of predicted values through successive recurrent computations, comparison of bidirectional predictions, and determination of substitute value data. Jung and Granson provide the outlier and extremeness considerations incorporated in the combined recurrent computations. The complete mapping of these limitations and the reason for combining the references are provided in the rejection of claim 1 and are incorporated herein. Accordingly, Fouladgar does not teach away from the proposed combination, and Applicant’s arguments concerning independent claim 1, 11, and 20 are not persuasive. Applicant further asserts that the dependent claims are patentable for the same reasons presented for the independent claims and requests individual consideration. Each dependent claim has been separately considered. The rejection identifies the additional limitation of each dependent claim and the corresponding disclosure in the cited reference. Applicant has no separately identified error in the rejection of any particular dependent claim. Because Applicant’s arguments concerning the independent claims are not persuasive, the assertion that the dependent claims are patentable for the same reasons is likewise not persuasive. Applicant’s arguments concerning claim 21 have been fully considered but are not persuasive. Claim 21 has been separately considered on its own merits. As set forth in the present rejection, Cao teaches recurrent neural network cells having a recurrent component implemented by a recurrent neural network and a regression component implemented by a fully connected network. These components correspond to the recurrent layer and regression layer recited in claim 21. Cao further performs the recurrent and regression operations in both the forward and backward directions. Accordingly, the additional limitation of claim 21 is addressed by Cao as set forth in the present rejection. For all the reason set forth above, the rejection under 35 U.S.C. 103 is maintained. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 5-9, 10-12, and 15-201are rejected under 35 U.S.C. 103 as being unpatentable over Fouladgar et al., (NPL: “A Novel LSTM for Multivariate Time Series with Massive Missingness” (Published: 2020)) in view of Chen at al., (Pub. No.: US 20120303411 A1 (Filed: 2011)) further in view of Granson et al., (Pat. No.: US 11,816,539 B1 (Filed: 2017)) further in view of Jung (Pub. No.: US 20220164688 A1 (Filed: 2020)) further in view of Cao et al., (NPL: “BRITS: Bidirectional Recurrent Imputation for Time Series” (Published: 2018)). Regarding claim 1, Fouladgar teaches the following: a system comprising: one or more processors, (Fouladgar, [page 14, section 5] “We proposed a novel LSTM-based model, FBVS-LSTM, consisting of four effective pieces of information as the augmentation of model input.”) a memory resource storing instructions, (Fouladgar, [page 14, section 5] “The results proved promising performance of the proposed model along with some LSTM-based derivative methods.”) that when executed by one or more processors, (Fouladgar, [page 14, section 5] “Finally, further evaluations with other deep-learning models like GRU will be performed…”) These statements make clear that the model was implemented, tested, and intended for use with deep learning architectures such as LSTMs and GRUs. It is well understood in the art that the implementation of LSTM-based models requires a computing system comprising one or more processors to execute model operations (e.g., matrix multiplications, forward/backward passes), and a memory resource to store model weights, input sequences and internal states. As such, it is inherent that the FBVS-LSTM system is implemented using one or more processors and memory resources storing instructions for model execution. However, Fouladgar does not teach but Fouladgar in view of Chen teaches: obtain a first time series data set corresponding to an aggregate retail order data set over a network, the first time series data set including a plurality of data elements, each data element including value data and corresponding time data; (Fouladgar, [page 4, section 3.2] Fouladgar describes multivariate time-series input as “X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.” Fouladgar teaches obtaining a time-series dataset including aggregated data. [Section 4.1] Specifically, Fouladgar describes the use of multivariate time-series inputs such as aggregated measurements of atmospheric and environmental variables, including “PM2.5 concentration, dew point, temperature, pressure, wind direction, cumulated wind speed, cumulated hours of snow and cumulated hours of rain.” [Abstract], Fouladgar also notes that the overall system receives “streaming data is collected by sensors or any other recording instruments.” Chen, [Abstract] “System, method and computer program product for demand modeling and prediction in retail categories. The method uses time-series data comprising of unit prices and unit sales for a designated choice set of related products, with the time-series data obtained over a given sequence of sales reporting periods, and over a collection of stores in a market geography. “– under the broadest reasonable interpretation, Fouladgar’s time-series dataset, composed of multivariate data arranged in temporal order, corresponds to the claimed time series dataset. Chen discloses retail order data arranged over time, corresponding the claimed aggregate “retail” order data. Additionally, because the data is collected from distributed sources, transmission of the time series values to the processing system is inherently over a network. Accordingly, Fouladgar in view of Chen teaches obtaining a first time-series dataset corresponding to an aggregate retail order data.) based on the first time series data set, (Fouladgar, [page 4, section 3.2] “time series X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.”) generate a second data set indicating one or more data elements of the plurality of data elements that are missing value data corresponding to missing aggregate retail order data (Fouladgar, [page 4, section 3.2] “Similar to [19], the missing data is formulated with a missing indicator M = {m1, m2,…,mT}T Є ℝ T×D for each observation x t d for each time series…Following the formulations, the missing indicator m at time stamp t of the variable d, m t d , is regarded as a binary mask…” – identifies which elements of the time-series dataset have missing value data, thereby generating the recited second data set. [Section 4.1] Fouladgar also teaches that the missing value data corresponds to missing aggregate order data because each time-series record contains aggregate of ordered measurements collected together at each time index, and the missing indicator identifies which values in the aggregate are missing. Fouladgar describes the aggregate ordered measurements at each timestamp including “PM2.5 concentration, dew point, temperature, pressure, wind direction, cumulated wind speed, cumulated hours of snow and cumulated hours of rain.” and further “PT08.S1 (tin oxide), PT08.S2 (titania), PT08.S3 (tungsten oxide), PT08.S4 (tungsten oxide) and PT08.S5 (indium oxide)” Chen, [Abstract] “System, method and computer program product for demand modeling and prediction in retail categories. The method uses time-series data comprising of unit prices and unit sales for a designated choice set of related products, with the time-series data obtained over a given sequence of sales reporting periods, and over a collection of stores in a market geography. “ – each time stamp forms aggregate order data record comprising these ordered measurements. Chen further teaches that such time-series data may comprise retail order data. When Fouladgar’s missing indicator M identifies a variable at a specific time as missing, it identifies missing value data within the aggregate order data record at each time step. Thus, under BRI, Fouladgar in view of Chen teaches generating a second data set indicating missing value data corresponding to missing aggregate retail order data.); execute at least one recurrent neural network to apply the first time series data set, (Fouladgar, [Abstract] “In this paper, we propose a novel model called forward and backward variable-sensitive LSTM (FBVS-LSTM) consisting of two decay mechanisms and some informative data.” [Page 2, Introduction] “Taking into account recurrent methods, the variants of Recurrent Neural Network (RNN) like Gated Recurrent Unit (GRU) [19–21] and Long Short-Term Memory (LSTM)” – and LSTM is a variant of an RNN. [page 4, section 3.2] “time series X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.”) the second data set (Fouladgar, [page 4, section 3.2] “Similar to [19], the missing data is formulated with a missing indicator M = {m1, m2,…,mT}T Є ℝ T×D for each observation x t d for each time series…Following the formulations, the missing indicator m at time stamp t of the variable d, m t d , is regarded as a binary mask…”) and the third data set, (Granson, [col., lines 7-17] “The imputation phase addresses potential issues related to sensitivity problems of neural networks to missing values with portions of a data set, and also by addressing potential issues caused by statistically significant outlier values. This poor quality data could otherwise reduce the effective size of the available training data set, which might otherwise weaken the predictive and classification strength of the overall mode. Imputation utilizes a series of replacement functions where missing or outlier interval variables are replaced by the median of non-missing and non-outlier values within the data set.”) implement a reconstruction operation on the at least one recurrent neural network to iteratively generate a substitute value data corresponding to substitute retail order data, for each data element of the one or more data elements that are missing value data. (Fouladgar, [page 6, section 3.3, Figure 1] “Since FB-LSTM only considers the missing indicator and the time intervals of missingness for each variable, we extend this model to a variable-sensitive version, namely forward and backward variable sensitive LSTM…Therefore, the missing rate of each variable is decayed within a negative exponential function to construct a missing factor β.” (β is computed to represent missingness) “Later, this factor directly takes part in the learning process in FBVS-LSTM.” (β is used in the model’s operations) “Incorporating the missing factor β in FB-LSTM, the gates rectified as the equations below and construct a final model as FBVS-LSTM.” (Gates are modified using β to produce outputs, including substitute values.) Fouladgar’s FBVS-LSTM architecture performs a reconstruction operation over missing data by encoding missingness (β), propagating through a recurrent neural network, and generating reconstructed outputs that serve as substituted values for each data element with missing value data. Under the broadest reasonable interpretation, such substitute values produced by the recurrent neural network constitute “substitute order data” because they are substituted (imputed) values within the same order time-series dataset. Chen, [Abstract] “System, method and computer program product for demand modeling and prediction in retail categories. The method uses time-series data comprising of unit prices and unit sales for a designated choice set of related products, with the time-series data obtained over a given sequence of sales reporting periods, and over a collection of stores in a market geography. “) transmit the first time series data and the substitute value data as a reconstructed data set over the network to train at least one machine learning model; and generate a retail order volume forecast via the at least one machine learning model (Fouladgar, [page 5, section 3.2] “To learn the parameters, the decay rates are imposed jointly to the input and hidden features of LSTM to capture the missing pattern informatively. This process constructs the main structure of LSTM with two decay mechanisms (FB-LSTM). In FB-LSTM, the missing data is imputed with the values either close to the mean of the variable or close to the last/first observation of the variable.” [section 4.1] “We reshape the samples and generate a multivariate time series of 24 h within 8, 5 and 6 variables, standing for each of our three datasets, respectively. These samples are required to feed into our models, discussed further in Section 4.3, for the purpose of short-term (next-hour) prediction.” Chen, paragraph [0100] “the data in the first 81 weeks (not shown) have been used for fitting the model parameters…The plot compares the predicted price dynamics 602 obtained from the method described herein with the actual price dynamics 605 that were observed during that time period…These results demonstrate the accuracy and reliability of the method for forecasting the future price dynamics from the historic sales data.” [0103] “FIG. 11 illustrates an example plot 850 of forecasted sales 855 obtained by combining the three method steps for obtaining the price prediction model, discrete choice model for market share, and time series model for market size…The sales between weeks 81 and 102 in FIG. 10 are forecasted using the method steps illustrated from FIG. 3 to FIG. 5, wherein the data in the first 81 weeks (not shown) are used for fitting model parameters.” – Fouladgar imputes missing values in multivariate time-series data and feeds the resulting time-series samples into an LSTM neural network to learn the model parameters. Under BRI, an LSTM neural network is a machine learning model. Fouladgar’s original and imputed values correspond to the first time-series data and substitute value data forming the reconstructed dataset, which is transmitted over the network as mapped above. Chen uses historical retail sales data to fit model parameters and uses the fitted models to forecast retail sales. In the combined system, Fouladgar’s imputation method is applied to Chen’s retail time-series data to produce a reconstructed dataset containing the original and substitute values. The reconstructed dataset is transmitted over the network to train the LSTM machine learning model, and the trained model is used to generate the claimed retail order volume forecast.) However, Fouladgar in view of Chen does not teach but, Fouladgar in view of Chen further in view Granson further in view of Jung teaches the limitation: based on the first time series data set , (Fouladgar, [page 4, section 3.2] “time series X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.”), iteratively implement a data imputation operation to determine a normality threshold based on a standard deviation of a mean value of the first time series (Jung, paragraph [0050] “The cluster outlier module (e.g., cluster outlier 212 and cluster outlier 222 of FIG. 2) can learn a column-wise model independently in parallel, where each column (i.e., sensor) can be modeled as a univariate Gaussian Mixture Model (GMM) with a number of K centroids. The probability density function of GMM with K centroids can be written by…Gaussian distribution of the random variable x with a mean   u k   and standard deviation   σ k of cluster k. An outlier can be defined by outlier clusters whose weight probability   w k is less than a user-defined threshold wmin and outlier points which are outside the confidence interval”) to generate a third data set including extremeness data indicating an extremeness score for each data element of the plurality of data elements; (Granson, [col. 14, lines 7-17] “The imputation phase addresses potential issues related to sensitivity problems of neural networks to missing values with portions of a data set, and also by addressing potential issues caused by statistically significant outlier values. This poor quality data could otherwise reduce the effective size of the available training data set, which might otherwise weaken the predictive and classification strength of the overall mode. Imputation utilizes a series of replacement functions where missing or outlier interval variables are replaced by the median of non-missing and non-outlier values within the data set.”) However, Fouladgar in view of Chen further in view of Granson further in view of Jung does not teach but Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao teaches the following limitations: wherein the at least one recurrent neural network includes: a first plurality of recurrent neural network cells arranged in a forward layer, wherein a first recurrent neural network cell of the first plurality of recurrent neural network cells provides a predicted output value and a predicted extremeness score as an input to a second recurrent neural network cell of the first plurality of recurrent neural network cells (Cao, [section 4.1] “In a unidirectional recurrent dynamical system, each value in the time series can be derived by its predecessors with a fixed arbitrary function [9, 24, 3]. Thus, we iteratively impute all the variables in the time series according to the recurrent dynamics.” [section 4.1.1] “We introduce a recurrent component and a regression component for imputation. The recurrent component is achieved by a recurrent neural network and the regression component is achieved by a fully-connected network…Eq. (1) is the regression component which transfers the hidden state h t - 1 to the estimated vector x ^ t . In Eq. (2), we replace missing values in x t with corresponding values in   t , and obtain the complement vector x t c …In Eq. (4), based on the decayed hidden state, we predict the next state   h t ” [section 4.2] “Formally, the BRITS-I algorithm performs the RITS-I as shown in Eq. (1) to Eq. (5) in forward and backward directions, respectively. In the forward direction, we obtain the estimation sequence { x ^ 1 ,   x ^ 2 , … , x ^ T } ” Jung, paragraph [0011] “ generating the cluster model based on the raw data and updating the cluster model based on the most recently imputed data or the filtered data comprises one or more of: determining, based on the raw data, the most recently imputed data, or the filtered data, clusters and information associated with the clusters, wherein the information associated with the clusters includes one or more of: a number of clusters; a centroid of a respective cluster; and a standard deviation associated with the respective cluster; classifying a cluster as an outlier cluster; classifying a point as an outlier point; and determining that the outlier point belongs to a first cluster of the determined clusters.” [0012] “generating, for a missing or null value based on a Gaussian distribution, a sample based on the determined clusters and the information associated with the clusters; and replacing the missing or null value with the generated sample.” [0013] “A probability density function of the GMM is based on a Gaussian distribution. An outlier cluster is defined based on a user-defined threshold. An outlier point is defined based on a user-defined confidence level.” Granson, [col. 15, lines 7-11] “The imputation phase addresses potential issues related to sensitivity problems of neural networks to missing values within portions of a data set, and also by addressing potential issues caused by statistically significant outlier values.” – Cao’s successive recurrent computations in the forward direction correspond to the claimed first plurality of recurrent neural network cells arranged in a forward layer. Cao generates estimated values, incorporates the estimated values into complement input, and uses the complement input to predict the recurrent state for the next recurrent computation. Thus, a predicted output value generated by a first forward recurrent cell is provided as input to a second forward recurrent cell. The corresponding predicted extremeness score is determined as mapped above through Jung and Granson and is incorporated with the predicted output value into a successive recurrent computation.); and a second plurality of recurrent neural network cells arranged in a backward layer, wherein a second recurrent neural network cell of the second plurality of recurrent neural network cells provides a predicted output value and a predicted extremeness score as an input to a first recurrent neural network cell of the second plurality of recurrent neural network cells (Cao, [section 4.2] “In this section, we propose an improved version called BRITS-I. The algorithm alleviates the above-mentioned issues by utilizing the bidirectional recurrent dynamics on the given time series, i.e., besides the forward direction, each value in time series can be also derived from the backward direction by another fixed arbitrary function…Consider the backward direction of the time series. In bidirectional recurrent dynamics, the estimation x ^ 4 reversely depends on x ^ 5 to x ^ 7 …Formally, the BRITS-I algorithm performs the RITS-I as shown in Eq. (1) to Eq. (5) in forward and backward directions, respectively…Similarly, in the backward direction, we obtain another estimation sequence { x ' ^ 1 ,   x ' ^ 2 , … , x ^ ' T } and another loss sequence { l 1 ' , l 2 ' … , l T ' } ” – Cao’s successive recurrent computations in the backwards direction correspond to the claimed second plurality of recurrent neural network cells arranged in a backward layer. Cao explains that an earlier version in the time series reversely depends on values occurring later in the time series and performs the same recurrent imputation operations in the backward direction. Thus, a predicted output value generated by a second backward recurrent cell is provided as input to a first backward recurrent cell. The corresponding predicted extremeness score is determined as mapped above through Jung and Granson and is incorporated with the predicted output value into the successive backward recurrent computation.), wherein a first predicted output value data and a first predicted extremeness score data of the forward layer is compared to a second predicted output value data and a second predicted extremeness score of the backward layer to identify the substitute value data (Cao, [section 4.2] “. In the forward direction, we obtain the estimation sequence { x ^ 1 ,   x ^ 2 , … , x ^ T } and the loss sequence { l 1 , l 2 , … , l T } . Similarly, in the backward direction, we obtain another estimation sequence { x ' ^ 1 ,   x ' ^ 2 , … , x ^ ' T } and another loss sequence { l 1 ' , l 2 ' … , l T ' } . We enforce the prediction in each step to be consistent in both directions by introducing the “consistency loss” … The final estimation in the t -th step is the mean of x ^ t , and x ' ^ t ” – Cao compares the predicted output values generated in the forward and backward layers using a consistency loss that evaluates the discrepancy between the bidirectional estimates. Cao then uses the mean of the forward and backward values to determine the final imputed value, corresponding to identifying the substitute-value data. In the combined system, the corresponding predicted extremeness scores mapped above through Jung and Granson are likewise compared for consistency between the forward and backward layers when identifying the substitute value data.); Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the teachings of Fouladgar, Granson, Jung, Chen, and Cao, wherein Fouladgar provides the time-series imputation framework using a recurrent neural network, Granson provides extremeness scoring and consideration of statistically significant outlier values during neural network imputation, Jung provides determining a normality threshold based on statistical measures including mean and standard deviation, Chen provides retail time-series data and retail forecasting, and Cao provides recurrent neural network cells operating in forward and backward layers, passing predicted values through successive recurrent computations, and comparing the forward and backward predictions to determine substitute value data. One would have been motivated to incorporate Granson’s extremeness scoring and Jung’s normality threshold determination into the recurrent imputation framework of Fouladgar, and to implement the framework using Cao’s bidirectional recurrent processing, so the predictions made from both temporal directions are evaluated for consistency and statistically extreme values are accounted for when determining substitute values. One would have further been motivated to apply the combined imputation technique to Chen’s retail time-series data to improve the accuracy and reliability of the reconstructed data used to train the model and generate the retail order volume forecast. Regarding claim 2, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view Cao teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1, Fouladgar further teaches: for each data element of the plurality of data elements that are missing value data, replace the data element with the corresponding substitute value data. (Fouladgar, [page 10, section 4.3] “We refer to these models as LSTM-0, LSTM-mean, B-LSTM, F-LSTM, and BVS-LSTM, respectively. The first model imputes missing data with zero…the third model imputes missingness considering the forward time interval…Finally, the last model employs both δ1 and µ, missing rate at each variable, for imputation.” This verifies that the models are created with substitute values for the missing data through imputation.) Regarding claim 5, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1, Fouladgar in view of Granson further in view of Jung further teaches: in a forward layer, determining a first set of predicted output values and a first set of predicted extremeness scores (Fouladgar, [page 6-7, section 3.3] “to make the model adapted to the missing rate of each variable independently, of contributing µd in the learning process of previously articulated model (FB-LSTM) and accordingly constructing FBVS-LSTM model…Therefore, the missing rate of each variable is decayed with a negative exponential function to construct a missing factor β. Later, the factor directly takes part in the learning process in FBVS-LSTM” (fist predicted extremeness scores) “Incorporating the missing factor β in FB-LSTM, the gates are rectified as the equations below and construct our final model as FBVS-LSTM.” (first predicted output values)) based on the first time series data set (Fouladgar, [page 4, section 3.2] “time series X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.”) and the third data set.( Granson, [col., lines 7-17] “The imputation phase addresses potential issues related to sensitivity problems of neural networks to missing values with portions of a data set, and also by addressing potential issues caused by statistically significant outlier values. This poor quality data could otherwise reduce the effective size of the available training data set, which might otherwise weaken the predictive and classification strength of the overall mode. Imputation utilizes a series of replacement functions where missing or outlier interval variables are replaced by the median of non-missing and non-outlier values within the data set.”) As explained in the analysis of claim 1, it would have been obvious to a person of ordinary skill in the art to substitute the decay value used in LSTM models with an extremeness score, as both serve to modulate or weight imputation based on characteristics of missing data. See discussion above. Regarding claim 6, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao teaches all the elements of claim 5, therefore is rejected for the same reasons as those presented for claim 5, Fouladgar in view of Granson further in view of Jung further teaches: wherein the plurality of data elements of the first time series data set includes at least a first data element and a second data element, and the extremeness data of the third data set includes at least a first extremeness value associated with the first data element and a second extremeness value associated with the second data element. (Fouladgar, [page 6-7, section 3.3] “to make the model adapted to the missing rate of each variable independently, there is a possibility of contributing µd in the learning process…” (multiple data elements (e.g., first and second)) “This missing rate of each variable is decayed with a negative exponential function to construct a missing factor β.” (The missing rate is computed per variable, and then transformed into β values (i.e., first extremeness value for the first data element, second for second, etc.) “Another parameter of FBVS-LSTM compared with FB-LSTM is vector P, responsible for learning the missing factor.” (The use of vector P confirms that this not a single score, but a set of learned values – meaning each data element has an associated extremeness-like value.)) As explained in the analysis of claim 1, it would have been obvious to a person of ordinary skill in the art to substitute the decay value used in LSTM models with an extremeness score, as both serve to modulate or weight imputation based on characteristics of missing data. See discussion above. Regarding claim 7, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao teaches all the elements of claim 6, therefore is rejected for the same reasons as those presented for claim 6, Fouladgar in view of Granson further in view of Jung further teaches: based on the first data element (Fouladgar, [page 7, section 3.3] “Incorporating the missing factor β in FB-LSTM, the gates are rectified as the equations below and construct our final model as FBVS-LSTM” This is used in combination with known input data at each time step (i.e., the first data element) as part of the computation.) and the first extremeness value (Fouladgar, [page 6, section 3.3] “The missing rate of each variable is decayed within a negative exponential function to construct a missing factor β. Later, this factor directly takes part in the learning process in FBVS-LSTM.” This shows β per variable and actively used.), determining a first predicted output value for at least one data element missing value data (Fouladgar, [page 4, section 3.1] “xt and ht-1 are the inputs and hidden state at time t and t-1” [page 7, section 3.3] “Incorporating the missing factor β in FB-LSTM, the gates are rectified as the equations below and constructed our final model as FBVS-LSTM.” Equations (20)-(23) define how β and the masked inputs (including missing values) influence internal gates. These gates compute the final output ht (predicted output value) via Equation (24).) and corresponding first predicted extremeness value for the at least one data element missing value data. (Fouladgar, [page 7, section 3.3] “Another parameter of FBVS-LSTM compared with FB-LSTM is the vector P, responsible for learning the missing factor.” This shows that the system learns and updates β using a parameter vector P – i.e. generates new extremeness values.) As explained in the analysis of claim 1, it would have been obvious to a person of ordinary skill in the art to substitute the decay value used in LSTM models with an extremeness score, as both serve to modulate or weight imputation based on characteristics of missing data. See discussion above. Regarding claim 8, Fouladgar in view of Chen further in view of Granson further in view of Jung teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7, Fouladgar in view of Granson further in view of Jung further teaches: based on the first predicted output value (Fouladgar, [page 7, section 3.3] “Incorporating the missing factor β in FB-LSTM, the gates are rectified as the equations below and constructed our final model as FBVS-LSTM.” Equations (20)-(24), where ht-1 is used in subsequent gate computations.) and the corresponding first predicted extremeness value (Fouladgar, [page 6, section 3.3] “The missing rate of each variable is decayed within a negative exponential function to construct a missing factor β. Later, this factor directly takes part in the learning process in FBVS-LSTM.), determining, a second predicted output value for at least another data element missing value data (Fouladgar, [page 4, section 3.2] “To deal with MNAR category of missing data, we modify LSTM with two decay mechanisms, namely forward and backward LSTM (FB-LSTM). The decay mechanisms are devised to reinforce the imputation of missing values more accurately. [page 5, section 3.2] “The Equation (11) shows the decay process over the input mathematically” This is where the model explicitly calculates the imputed value for missing input at time t step.) and corresponding second predicted extremeness value (Fouladgar, [page 7, section 3.3] “Another parameter of FBVS-LSTM compared with FB-LSTM is the vector P, responsible for learning the missing factor.” This shows that the system learns and updates β using a parameter vector P – i.e. generates new extremeness values.) As explained in the analysis of claim 1, it would have been obvious to a person of ordinary skill in the art to substitute the decay value used in LSTM models with an extremeness score, as both serve to modulate or weight imputation based on characteristics of missing data. See discussion above. Regarding claim 9, Fouladgar in view of Chen further in view of Granson further in view of Jung teaches all the elements of claim 5, therefore is rejected for the same reasons as those presented for claim 5, Fouladgar in view of Granson further in view of Jung further teaches: in a backward layer, (Fouladgar, [page 6, section 3.3] “we extend this model to a variable-sensitive LSTM (FBVS-LSTM).”) determining a second set of predicted output values (Fouladgar, page 7, section 3.3] “the gates are rectified as equations below construct our final model as FBVS-LSTM.” Followed by Equations (20) - (24) where ht is the output at each step.) and a second set of predicted extremeness scores (Fouladgar, [page 6, section 3.3] “The missing rate of each variable is decayed within a negative exponential function to construct a missing factor β. Later, this factor directly takes part in the learning process in FBVS-LSTM.” This shows β per variable and actively used.), based on the first time series data set (Fouladgar, [page 4, section 3.2] “time series X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.”) and the third data set (Granson, [col. 14, lines 7-17] “The imputation phase addresses potential issues related to sensitivity problems of neural networks to missing values with portions of a data set, and also by addressing potential issues caused by statistically significant outlier values. This poor quality data could otherwise reduce the effective size of the available training data set, which might otherwise weaken the predictive and classification strength of the overall mode. Imputation utilizes a series of replacement functions where missing or outlier interval variables are replaced by the median of non-missing and non-outlier values within the data set.”). As explained in the analysis of claim 1, it would have been obvious to a person of ordinary skill in the art to substitute the decay value used in LSTM models with an extremeness score, as both serve to modulate or weight imputation based on characteristics of missing data. See discussion above. Regarding claim 10, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao teaches all the elements of claim 5, therefore is rejected for the same reasons as those presented for claim 5, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao teaches: based at least on the first set of predicted output values, (Fouladgar, [page 7, section 3.3] “Incorporating the missing factor β in FB-LSTM, the gates are rectified as the equations below and constructed our final model as FBVS-LSTM.” Equations (20)-(24), where ht-1 is used in subsequent gate computations.) the first set of predicted extremeness scores, (Fouladgar, [page 6, section 3.3] “The missing rate of each variable is decayed within a negative exponential function to construct a missing factor β. Later, this factor directly takes part in the learning process in FBVS-LSTM.” This shows β per variable and actively used.) the second set of predicted output values (Fouladgar, page 7, section 3.3] “the gates are rectified as equations below construct our final model as FBVS-LSTM.” Followed by Equations (20) - (24) where ht is the output at each step.) and the second set of predicted extremeness scores, (Fouladgar, [page 6, section 3.3] “The missing rate of each variable is decayed within a negative exponential function to construct a missing factor β. Later, this factor directly takes part in the learning process in FBVS-LSTM.” This shows β per variable and actively used.), generating discrepancy data (Cao, [page 5, section 4.2] “We enforce the prediction in each step to be consistent in both direction by introducing the ‘consistency loss’…where we use the mean squared error as the discrepancy in our experiment. The final estimation loss is obtained by accumulating the forward loss, the backward loss, and the consistency loss.”) Regarding claim 11, Fouladgar in view of Chen teaches the following: obtain a first time series data set corresponding to an aggregate retail order data set over a network, the first time series data set including a plurality of data elements, each data element including value data and corresponding time data; (Fouladgar, [page 4, section 3.2] Fouladgar describes multivariate time-series input as “X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.” Fouladgar teaches obtaining a time-series dataset including aggregated data. [Section 4.1] Specifically, Fouladgar describes the use of multivariate time-series inputs such as aggregated measurements of atmospheric and environmental variables, including “PM2.5 concentration, dew point, temperature, pressure, wind direction, cumulated wind speed, cumulated hours of snow and cumulated hours of rain.” [Abstract], Fouladgar also notes that the overall system receives “streaming data is collected by sensors or any other recording instruments.” Chen, [Abstract] “System, method and computer program product for demand modeling and prediction in retail categories. The method uses time-series data comprising of unit prices and unit sales for a designated choice set of related products, with the time-series data obtained over a given sequence of sales reporting periods, and over a collection of stores in a market geography. “ – under the broadest reasonable interpretation, Fouladgar’s time-series dataset, composed of multivariate data arranged in temporal order, corresponds to the claimed time series dataset. Chen discloses retail order data arranged over time, corresponding the claimed aggregate “retail” order data. Additionally, because the data is collected from distributed sources, transmission of the time series values to the processing system is inherently over a network. Accordingly, Fouladgar in view of Chen teaches obtaining a first time-series dataset corresponding to an aggregate retail order data.) based on the first time series data set, (Fouladgar, [page 4, section 3.2] “time series X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.”) generate a second data set indicating one or more data elements of the plurality of data elements that are missing value data corresponding to missing aggregate retail order data (Fouladgar, [page 4, section 3.2] “Similar to [19], the missing data is formulated with a missing indicator M = {m1, m2,…,mT}T Є ℝ T×D for each observation x t d for each time series…Following the formulations, the missing indicator m at time stamp t of the variable d, m t d , is regarded as a binary mask…” – identifies which elements of the time-series dataset have missing value data, thereby generating the recited second data set. [Section 4.1] Fouladgar also teaches that the missing value data corresponds to missing aggregate order data because each time-series record contains aggregate of ordered measurements collected together at each time index, and the missing indicator identifies which values in the aggregate are missing. Fouladgar describes the aggregate ordered measurements at each timestamp including “PM2.5 concentration, dew point, temperature, pressure, wind direction, cumulated wind speed, cumulated hours of snow and cumulated hours of rain.” and further “PT08.S1 (tin oxide), PT08.S2 (titania), PT08.S3 (tungsten oxide), PT08.S4 (tungsten oxide) and PT08.S5 (indium oxide)” Chen, [Abstract] “System, method and computer program product for demand modeling and prediction in retail categories. The method uses time-series data comprising of unit prices and unit sales for a designated choice set of related products, with the time-series data obtained over a given sequence of sales reporting periods, and over a collection of stores in a market geography. “– each time stamp forms aggregate order data record comprising these ordered measurements. Chen further teaches that such time-series data may comprise retail order data. When Fouladgar’s missing indicator M identifies a variable at a specific time as missing, it identifies missing value data within the aggregate order data record at each time step. Thus, under BRI, Fouladgar in view of Chen teaches generating a second data set indicating missing value data corresponding to missing aggregate retail order data.); execute at least one recurrent neural network to apply the first time series data set, (Fouladgar, [Abstract] “In this paper, we propose a novel model called forward and backward variable-sensitive LSTM (FBVS-LSTM) consisting of two decay mechanisms and some informative data.” [Page 2, Introduction] “Taking into account recurrent methods, the variants of Recurrent Neural Network (RNN) like Gated Recurrent Unit (GRU) [19–21] and Long Short-Term Memory (LSTM)” – and LSTM is a variant of an RNN. [page 4, section 3.2] “time series X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.”) the second data set (Fouladgar, [page 4, section 3.2] “Similar to [19], the missing data is formulated with a missing indicator M = {m1, m2,…,mT}T Є ℝ T×D for each observation x t d for each time series…Following the formulations, the missing indicator m at time stamp t of the variable d, m t d , is regarded as a binary mask…”) and the third data set, (Granson, [col., lines 7-17] “The imputation phase addresses potential issues related to sensitivity problems of neural networks to missing values with portions of a data set, and also by addressing potential issues caused by statistically significant outlier values. This poor quality data could otherwise reduce the effective size of the available training data set, which might otherwise weaken the predictive and classification strength of the overall mode. Imputation utilizes a series of replacement functions where missing or outlier interval variables are replaced by the median of non-missing and non-outlier values within the data set.”) implement a reconstruction operation on the at least one recurrent neural network to iteratively generate a substitute value data corresponding to substitute retail order data, for each data element of the one or more data elements that are missing value data. (Fouladgar, [page 6, section 3.3, Figure 1] “Since FB-LSTM only considers the missing indicator and the time intervals of missingness for each variable, we extend this model to a variable-sensitive version, namely forward and backward variable sensitive LSTM…Therefore, the missing rate of each variable is decayed within a negative exponential function to construct a missing factor β.” (β is computed to represent missingness) “Later, this factor directly takes part in the learning process in FBVS-LSTM.” (β is used in the model’s operations) “Incorporating the missing factor β in FB-LSTM, the gates rectified as the equations below and construct a final model as FBVS-LSTM.” (Gates are modified using β to produce outputs, including substitute values.) Fouladgar’s FBVS-LSTM architecture performs a reconstruction operation over missing data by encoding missingness (β), propagating through a recurrent neural network, and generating reconstructed outputs that serve as substituted values for each data element with missing value data. Under the broadest reasonable interpretation, such substitute values produced by the recurrent neural network constitute “substitute order data” because they are substituted (imputed) values within the same order time-series dataset. Chen, [Abstract] “System, method and computer program product for demand modeling and prediction in retail categories. The method uses time-series data comprising of unit prices and unit sales for a designated choice set of related products, with the time-series data obtained over a given sequence of sales reporting periods, and over a collection of stores in a market geography.) transmitting the first time series data and the substitute value data as a reconstructed data over the network to train at least one machine learning model; and generating an order volume forecast (Fouladgar, [page 5, section 3.2] “To learn the parameters, the decay rates are imposed jointly to the input and hidden features of LSTM to capture the missing pattern informatively. This process constructs the main structure of LSTM with two decay mechanisms (FB-LSTM). In FB-LSTM, the missing data is imputed with the values either close to the mean of the variable or close to the last/first observation of the variable.” [section 4.1] “We reshape the samples and generate a multivariate time series of 24 h within 8, 5 and 6 variables, standing for each of our three datasets, respectively. These samples are required to feed into our models, discussed further in Section 4.3, for the purpose of short-term (next-hour) prediction.” Chen, paragraph [0100] “the data in the first 81 weeks (not shown) have been used for fitting the model parameters…The plot compares the predicted price dynamics 602 obtained from the method described herein with the actual price dynamics 605 that were observed during that time period…These results demonstrate the accuracy and reliability of the method for forecasting the future price dynamics from the historic sales data.” [0103] “FIG. 11 illustrates an example plot 850 of forecasted sales 855 obtained by combining the three method steps for obtaining the price prediction model, discrete choice model for market share, and time series model for market size…The sales between weeks 81 and 102 in FIG. 10 are forecasted using the method steps illustrated from FIG. 3 to FIG. 5, wherein the data in the first 81 weeks (not shown) are used for fitting model parameters.” – Fouladgar imputes missing values in multivariate time-series data and feeds the resulting time-series samples into an LSTM neural network to learn the model parameters. Under BRI, an LSTM neural network is a machine learning model. Fouladgar’s original and imputed values correspond to the first time-series data and substitute value data forming the reconstructed dataset, which is transmitted over the network as mapped above. Chen uses historical retail sales data to fit model parameters and uses the fitted models to forecast retail sales. In the combined system, Fouladgar’s imputation method is applied to Chen’s retail time-series data to produce a reconstructed dataset containing the original and substitute values. The reconstructed dataset is transmitted over the network to train the LSTM machine learning model, and the trained model is used to generate the claimed retail order volume forecast.). However, Fouladgar in view of Chen does not teach but, Fouladgar in view of Chen further in view of Granson further in view of Jung teaches the limitation: based on the first time series data set , (Fouladgar, [page 4, section 3.2] “time series X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.”), iteratively implement a data imputation operation to determine a normality threshold based on a standard deviation of a mean value of the first time series (Jung, paragraph [0050] “The cluster outlier module (e.g., cluster outlier 212 and cluster outlier 222 of FIG. 2) can learn a column-wise model independently in parallel, where each column (i.e., sensor) can be modeled as a univariate Gaussian Mixture Model (GMM) with a number of K centroids. The probability density function of GMM with K centroids can be written by…Gaussian distribution of the random variable x with a mean   u k   and standard deviation   σ k of cluster k. An outlier can be defined by outlier clusters whose weight probability   w k is less than a user-defined threshold wmin and outlier points which are outside the confidence interval”) to generate a third data set including extremeness data indicating an extremeness score for each data element of the plurality of data elements; (Granson, [col. 14, lines 7-17] “The imputation phase addresses potential issues related to sensitivity problems of neural networks to missing values with portions of a data set, and also by addressing potential issues caused by statistically significant outlier values. This poor quality data could otherwise reduce the effective size of the available training data set, which might otherwise weaken the predictive and classification strength of the overall mode. Imputation utilizes a series of replacement functions where missing or outlier interval variables are replaced by the median of non-missing and non-outlier values within the data set.”) However, Fouladgar in view of Chen further in view of Granson further in view of Jung does not teach but Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao teaches the following limitations: wherein the at least one recurrent neural network includes: a first plurality of recurrent neural network cells arranged in a forward layer, wherein a first recurrent neural network cell of the first plurality of recurrent neural network cells provides a predicted output value and a predicted extremeness score as an input to a second recurrent neural network cell of the first plurality of recurrent neural network cells (Cao, [section 4.1] “In a unidirectional recurrent dynamical system, each value in the time series can be derived by its predecessors with a fixed arbitrary function [9, 24, 3]. Thus, we iteratively impute all the variables in the time series according to the recurrent dynamics.” [section 4.1.1] “We introduce a recurrent component and a regression component for imputation. The recurrent component is achieved by a recurrent neural network and the regression component is achieved by a fully-connected network…Eq. (1) is the regression component which transfers the hidden state h t - 1 to the estimated vector x ^ t . In Eq. (2), we replace missing values in x t with corresponding values in   t , and obtain the complement vector x t c …In Eq. (4), based on the decayed hidden state, we predict the next state   h t ” [section 4.2] “Formally, the BRITS-I algorithm performs the RITS-I as shown in Eq. (1) to Eq. (5) in forward and backward directions, respectively. In the forward direction, we obtain the estimation sequence { x ^ 1 ,   x ^ 2 , … , x ^ T } ” Jung, paragraph [0011] “ generating the cluster model based on the raw data and updating the cluster model based on the most recently imputed data or the filtered data comprises one or more of: determining, based on the raw data, the most recently imputed data, or the filtered data, clusters and information associated with the clusters, wherein the information associated with the clusters includes one or more of: a number of clusters; a centroid of a respective cluster; and a standard deviation associated with the respective cluster; classifying a cluster as an outlier cluster; classifying a point as an outlier point; and determining that the outlier point belongs to a first cluster of the determined clusters.” [0012] “generating, for a missing or null value based on a Gaussian distribution, a sample based on the determined clusters and the information associated with the clusters; and replacing the missing or null value with the generated sample.” [0013] “A probability density function of the GMM is based on a Gaussian distribution. An outlier cluster is defined based on a user-defined threshold. An outlier point is defined based on a user-defined confidence level.” Granson, [col. 15, lines 7-11] “The imputation phase addresses potential issues related to sensitivity problems of neural networks to missing values within portions of a data set, and also by addressing potential issues caused by statistically significant outlier values.” – Cao’s successive recurrent computations in the forward direction correspond to the claimed first plurality of recurrent neural network cells arranged in a forward layer. Cao generates estimated values, incorporates the estimated values into complement input, and uses the complement input to predict the recurrent state for the next recurrent computation. Thus, a predicted output value generated by a first forward recurrent cell is provided as input to a second forward recurrent cell. The corresponding predicted extremeness score is determined as mapped above through Jung and Granson and is incorporated with the predicted output value into a successive recurrent computation.); and a second plurality of recurrent neural network cells arranged in a backward layer, wherein a second recurrent neural network cell of the second plurality of recurrent neural network cells provides a predicted output value and a predicted extremeness score as an input to a first recurrent neural network cell of the second plurality of recurrent neural network cells (Cao, [section 4.2] “In this section, we propose an improved version called BRITS-I. The algorithm alleviates the above-mentioned issues by utilizing the bidirectional recurrent dynamics on the given time series, i.e., besides the forward direction, each value in time series can be also derived from the backward direction by another fixed arbitrary function…Consider the backward direction of the time series. In bidirectional recurrent dynamics, the estimation x ^ 4 reversely depends on x ^ 5 to x ^ 7 …Formally, the BRITS-I algorithm performs the RITS-I as shown in Eq. (1) to Eq. (5) in forward and backward directions, respectively…Similarly, in the backward direction, we obtain another estimation sequence { x ' ^ 1 ,   x ' ^ 2 , … , x ^ ' T } and another loss sequence { l 1 ' , l 2 ' … , l T ' } ” – Cao’s successive recurrent computations in the backwards direction correspond to the claimed second plurality of recurrent neural network cells arranged in a backward layer. Cao explains that an earlier version in the time series reversely depends on values occurring later in the time series and performs the same recurrent imputation operations in the backward direction. Thus, a predicted output value generated by a second backward recurrent cell is provided as input to a first backward recurrent cell. The corresponding predicted extremeness score is determined as mapped above through Jung and Granson and is incorporated with the predicted output value into the successive backward recurrent computation.), wherein a first predicted output value data and a first predicted extremeness score data of the forward layer is compared to a second predicted output value data and a second predicted extremeness score of the backward layer to identify the substitute value data (Cao, [section 4.2] “. In the forward direction, we obtain the estimation sequence { x ^ 1 ,   x ^ 2 , … , x ^ T } and the loss sequence { l 1 , l 2 , … , l T } . Similarly, in the backward direction, we obtain another estimation sequence { x ' ^ 1 ,   x ' ^ 2 , … , x ^ ' T } and another loss sequence { l 1 ' , l 2 ' … , l T ' } . We enforce the prediction in each step to be consistent in both directions by introducing the “consistency loss” … The final estimation in the t -th step is the mean of x ^ t , and x ' ^ t ” – Cao compares the predicted output values generated in the forward and backward layers using a consistency loss that evaluates the discrepancy between the bidirectional estimates. Cao then uses the mean of the forward and backward values to determine the final imputed value, corresponding to identifying the substitute-value data. In the combined system, the corresponding predicted extremeness scores mapped above through Jung and Granson are likewise compared for consistency between the forward and backward layers when identifying the substitute value data.); Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the teachings of Fouladgar, Granson, Jung, Chen, and Cao, wherein Fouladgar provides the time-series imputation framework using a recurrent neural network, Granson provides extremeness scoring and consideration of statistically significant outlier values during neural network imputation, Jung provides determining a normality threshold based on statistical measures including mean and standard deviation, Chen provides retail time-series data and retail forecasting, and Cao provides recurrent neural network cells operating in forward and backward layers, passing predicted values through successive recurrent computations, and comparing the forward and backward predictions to determine substitute value data. One would have been motivated to incorporate Granson’s extremeness scoring and Jung’s normality threshold determination into the recurrent imputation framework of Fouladgar, and to implement the framework using Cao’s bidirectional recurrent processing, so the predictions made from both temporal directions are evaluated for consistency and statistically extreme values are accounted for when determining substitute values. One would have further been motivated to apply the combined imputation technique to Chen’s retail time-series data to improve the accuracy and reliability of the reconstructed data used to train the model and generate the retail order volume forecast. Regarding claim 12, Fouladgar in view of Chen further in view of Granson further in view of Jung teaches all the elements of claim 11, therefore is rejected for the same reasons as those presented for claim 11, the claim recites similar limitations corresponding to claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale. Regarding claim 15, Fouladgar in view of Chen further in view of Granson further in view of Jung teaches all the elements of claim 11, therefore is rejected for the same reasons as those presented for claim 11, the claim recites similar limitations corresponding to claim 5 and is rejected for similar reasons as claim 5 using similar teachings and rationale. Regarding claim 16, Fouladgar in view of Chen further in view of Granson further in view of Jung teaches all the elements of claim 15, therefore is rejected for the same reasons as those presented for claim 15, the claim recites similar limitations corresponding to claim 6 and is rejected for similar reasons as claim 6 using similar teachings and rationale. Regarding claim 17, Fouladgar in view of Chen further in view of Granson further in view of Jung teaches all the elements of claim 16, therefore is rejected for the same reasons as those presented for claim 16, the claim recites similar limitations corresponding to claim 7 and is rejected for similar reasons as claim 7 using similar teachings and rationale. Regarding claim 18, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao teaches all the elements of claim 17, therefore is rejected for the same reasons as those presented for claim 17. The claim recites similar limitations corresponding to claim 8 and is rejected for similar reasons as claim 8 using similar teachings and rationale. Regarding claim 19, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao teaches all the elements of claim 15, therefore is rejected for the same reasons as those presented for claim 15. The claim recites similar limitations corresponding to claims 9 and 10. Therefore, the claim is rejected for similar reasons as claims 9 and 10 using similar teachings and rationale. Regarding claim 20, Fouladgar teaches the following: A non-transitory computer readable medium having instructions stored thereon, (Fouladgar, [page 14, section 5] “We proposed a novel LSTM-based model, FBVS-LSTM, consisting of four effective pieces of information as the augmentation of model input.”) wherein the instructions, when executed by one or more processors, (Fouladgar, [page 14, section 5] “Future study aims to replicate the experiments in health-related domain, mainly a daily stress-monitoring dataset collected from different sensors…Finally, further evaluations with deep-learning models like GRU will be performed…”)cause a computing device to: These statements indicate that the model was developed, tested, and deployed for real-world use across datasets, including sensor-based experiments and deep learning evaluation. It is well understood in the art that such LSTM-based models are implemented as computer-executable code and require a computing device to operate. This ability to replicate experiments and perform model evaluations necessarily implies execution of stored instructions on a computing system. As such, it is inherent that the FBVS-LSTM system includes a non-transitory computer readable medium storing instruction, and that those instructions are executed by one or more processors to perform model operations such as training, evaluation, and inference. However, Fouladgar does not teach but Fouladgar in view of Chen teaches the following limitations: obtain a first time series data set corresponding to an aggregate retail order data set over a network, the first time series data set including a plurality of data elements, each data element including value data and corresponding time data; (Fouladgar, [page 4, section 3.2] Fouladgar describes multivariate time-series input as “X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.” Fouladgar teaches obtaining a time-series dataset including aggregated data. [Section 4.1] Specifically, Fouladgar describes the use of multivariate time-series inputs such as aggregated measurements of atmospheric and environmental variables, including “PM2.5 concentration, dew point, temperature, pressure, wind direction, cumulated wind speed, cumulated hours of snow and cumulated hours of rain.” [Abstract], Fouladgar also notes that the overall system receives “streaming data is collected by sensors or any other recording instruments.” Chen, [Abstract] “System, method and computer program product for demand modeling and prediction in retail categories. The method uses time-series data comprising of unit prices and unit sales for a designated choice set of related products, with the time-series data obtained over a given sequence of sales reporting periods, and over a collection of stores in a market geography. “– under the broadest reasonable interpretation, Fouladgar’s time-series dataset, composed of multivariate data arranged in temporal order, corresponds to the claimed time series dataset. Chen discloses retail order data arranged over time, corresponding the claimed aggregate “retail” order data. Additionally, because the data is collected from distributed sources, transmission of the time series values to the processing system is inherently over a network. Accordingly, Fouladgar in view of Chen teaches obtaining a first time-series dataset corresponding to an aggregate retail order data.) based on the first time series data set, (Fouladgar, [page 4, section 3.2] “time series X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.”) generate a second data set indicating one or more data elements of the plurality of data elements that are missing value data corresponding to missing aggregate retail order data (Fouladgar, [page 4, section 3.2] “Similar to [19], the missing data is formulated with a missing indicator M = {m1, m2,…,mT}T Є ℝ T×D for each observation x t d for each time series…Following the formulations, the missing indicator m at time stamp t of the variable d, m t d , is regarded as a binary mask…” – identifies which elements of the time-series dataset have missing value data, thereby generating the recited second data set. [Section 4.1] Fouladgar also teaches that the missing value data corresponds to missing aggregate order data because each time-series record contains aggregate of ordered measurements collected together at each time index, and the missing indicator identifies which values in the aggregate are missing. Fouladgar describes the aggregate ordered measurements at each timestamp including “PM2.5 concentration, dew point, temperature, pressure, wind direction, cumulated wind speed, cumulated hours of snow and cumulated hours of rain.” and further “PT08.S1 (tin oxide), PT08.S2 (titania), PT08.S3 (tungsten oxide), PT08.S4 (tungsten oxide) and PT08.S5 (indium oxide)” Chen [Abstract] “System, method and computer program product for demand modeling and prediction in retail categories. The method uses time-series data comprising of unit prices and unit sales for a designated choice set of related products, with the time-series data obtained over a given sequence of sales reporting periods, and over a collection of stores in a market geography.” – each time stamp forms aggregate order data record comprising these ordered measurements. Chen further teaches that such time-series data may comprise retail order data. When Fouladgar’s missing indicator M identifies a variable at a specific time as missing, it identifies missing value data within the aggregate order data record at each time step. Thus, under BRI, Fouladgar in view of Chen teaches generating a second data set indicating missing value data corresponding to missing aggregate retail order data.); execute at least one recurrent neural network to apply the first time series data set, (Fouladgar, [Abstract] “In this paper, we propose a novel model called forward and backward variable-sensitive LSTM (FBVS-LSTM) consisting of two decay mechanisms and some informative data.” [Page 2, Introduction] “Taking into account recurrent methods, the variants of Recurrent Neural Network (RNN) like Gated Recurrent Unit (GRU) [19–21] and Long Short-Term Memory (LSTM)” – and LSTM is a variant of an RNN. [page 4, section 3.2] “time series X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.”) the second data set (Fouladgar, [page 4, section 3.2] “Similar to [19], the missing data is formulated with a missing indicator M = {m1, m2,…,mT}T Є ℝ T×D for each observation x t d for each time series…Following the formulations, the missing indicator m at time stamp t of the variable d, m t d , is regarded as a binary mask…”) and the third data set, (Granson, [col., lines 7-17] “The imputation phase addresses potential issues related to sensitivity problems of neural networks to missing values with portions of a data set, and also by addressing potential issues caused by statistically significant outlier values. This poor quality data could otherwise reduce the effective size of the available training data set, which might otherwise weaken the predictive and classification strength of the overall mode. Imputation utilizes a series of replacement functions where missing or outlier interval variables are replaced by the median of non-missing and non-outlier values within the data set.”) implement a reconstruction operation on the at least one recurrent neural network to iteratively generate a substitute value data corresponding to substitute retail order data, for each data element of the one or more data elements that are missing value data. (Fouladgar, [page 6, section 3.3, Figure 1] “Since FB-LSTM only considers the missing indicator and the time intervals of missingness for each variable, we extend this model to a variable-sensitive version, namely forward and backward variable sensitive LSTM…Therefore, the missing rate of each variable is decayed within a negative exponential function to construct a missing factor β.” (β is computed to represent missingness) “Later, this factor directly takes part in the learning process in FBVS-LSTM.” (β is used in the model’s operations) “Incorporating the missing factor β in FB-LSTM, the gates rectified as the equations below and construct a final model as FBVS-LSTM.” (Gates are modified using β to produce outputs, including substitute values.) Fouladgar’s FBVS-LSTM architecture performs a reconstruction operation over missing data by encoding missingness (β), propagating through a recurrent neural network, and generating reconstructed outputs that serve as substituted values for each data element with missing value data. Under the broadest reasonable interpretation, such substitute values produced by the recurrent neural network constitute “substitute order data” because they are substituted (imputed) values within the same order time-series dataset. Chen, [Abstract] “System, method and computer program product for demand modeling and prediction in retail categories. The method uses time-series data comprising of unit prices and unit sales for a designated choice set of related products, with the time-series data obtained over a given sequence of sales reporting periods, and over a collection of stores in a market geography.) transmitting the first time series data and the substitute value data as reconstructed data set over the network to train at least one machine learning model; and generating an order volume forecast (Fouladgar, [page 5, section 3.2] “To learn the parameters, the decay rates are imposed jointly to the input and hidden features of LSTM to capture the missing pattern informatively. This process constructs the main structure of LSTM with two decay mechanisms (FB-LSTM). In FB-LSTM, the missing data is imputed with the values either close to the mean of the variable or close to the last/first observation of the variable.” [section 4.1] “We reshape the samples and generate a multivariate time series of 24 h within 8, 5 and 6 variables, standing for each of our three datasets, respectively. These samples are required to feed into our models, discussed further in Section 4.3, for the purpose of short-term (next-hour) prediction.” Chen, paragraph [0100] “the data in the first 81 weeks (not shown) have been used for fitting the model parameters…The plot compares the predicted price dynamics 602 obtained from the method described herein with the actual price dynamics 605 that were observed during that time period…These results demonstrate the accuracy and reliability of the method for forecasting the future price dynamics from the historic sales data.” [0103] “FIG. 11 illustrates an example plot 850 of forecasted sales 855 obtained by combining the three method steps for obtaining the price prediction model, discrete choice model for market share, and time series model for market size…The sales between weeks 81 and 102 in FIG. 10 are forecasted using the method steps illustrated from FIG. 3 to FIG. 5, wherein the data in the first 81 weeks (not shown) are used for fitting model parameters.” – Fouladgar imputes missing values in multivariate time-series data and feeds the resulting time-series samples into an LSTM neural network to learn the model parameters. Under BRI, an LSTM neural network is a machine learning model. Fouladgar’s original and imputed values correspond to the first time-series data and substitute value data forming the reconstructed dataset, which is transmitted over the network as mapped above. Chen uses historical retail sales data to fit model parameters and uses the fitted models to forecast retail sales. In the combined system, Fouladgar’s imputation method is applied to Chen’s retail time-series data to produce a reconstructed dataset containing the original and substitute values. The reconstructed dataset is transmitted over the network to train the LSTM machine learning model, and the trained model is used to generate the claimed retail order volume forecast.). However, Fouladgar in view of Chen does not teach but, Fouladgar in view of Chen further in view of Granson further in view of Jung teaches the limitation: based on the first time series data set , (Fouladgar, [page 4, section 3.2] “time series X = {x1, x2,…,xT}T Є ℝT×D where d Є {1,…,D} and t Є {1,…,T} denote the variable and time of observation, respectively.”), iteratively implement a data imputation operation to determine a normality threshold based on a standard deviation of a mean value of the first time series (Jung, paragraph [0050] “The cluster outlier module (e.g., cluster outlier 212 and cluster outlier 222 of FIG. 2) can learn a column-wise model independently in parallel, where each column (i.e., sensor) can be modeled as a univariate Gaussian Mixture Model (GMM) with a number of K centroids. The probability density function of GMM with K centroids can be written by…Gaussian distribution of the random variable x with a mean   u k   and standard deviation   σ k of cluster k. An outlier can be defined by outlier clusters whose weight probability   w k is less than a user-defined threshold wmin and outlier points which are outside the confidence interval”) to generate a third data set including extremeness data indicating an extremeness score for each data element of the plurality of data elements; (Granson, [col. 14, lines 7-17] “The imputation phase addresses potential issues related to sensitivity problems of neural networks to missing values with portions of a data set, and also by addressing potential issues caused by statistically significant outlier values. This poor quality data could otherwise reduce the effective size of the available training data set, which might otherwise weaken the predictive and classification strength of the overall mode. Imputation utilizes a series of replacement functions where missing or outlier interval variables are replaced by the median of non-missing and non-outlier values within the data set.”) However, Fouladgar in view of Chen further in view of Granson further in view of Jung does not teach but Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao teaches the following limitations: wherein the at least one recurrent neural network includes: a first plurality of recurrent neural network cells arranged in a forward layer, wherein a first recurrent neural network cell of the first plurality of recurrent neural network cells provides a predicted output value and a predicted extremeness score as an input to a second recurrent neural network cell of the first plurality of recurrent neural network cells (Cao, [section 4.1] “In a unidirectional recurrent dynamical system, each value in the time series can be derived by its predecessors with a fixed arbitrary function [9, 24, 3]. Thus, we iteratively impute all the variables in the time series according to the recurrent dynamics.” [section 4.1.1] “We introduce a recurrent component and a regression component for imputation. The recurrent component is achieved by a recurrent neural network and the regression component is achieved by a fully-connected network…Eq. (1) is the regression component which transfers the hidden state h t - 1 to the estimated vector x ^ t . In Eq. (2), we replace missing values in x t with corresponding values in   t , and obtain the complement vector x t c …In Eq. (4), based on the decayed hidden state, we predict the next state   h t ” [section 4.2] “Formally, the BRITS-I algorithm performs the RITS-I as shown in Eq. (1) to Eq. (5) in forward and backward directions, respectively. In the forward direction, we obtain the estimation sequence { x ^ 1 ,   x ^ 2 , … , x ^ T } ” Jung, paragraph [0011] “ generating the cluster model based on the raw data and updating the cluster model based on the most recently imputed data or the filtered data comprises one or more of: determining, based on the raw data, the most recently imputed data, or the filtered data, clusters and information associated with the clusters, wherein the information associated with the clusters includes one or more of: a number of clusters; a centroid of a respective cluster; and a standard deviation associated with the respective cluster; classifying a cluster as an outlier cluster; classifying a point as an outlier point; and determining that the outlier point belongs to a first cluster of the determined clusters.” [0012] “generating, for a missing or null value based on a Gaussian distribution, a sample based on the determined clusters and the information associated with the clusters; and replacing the missing or null value with the generated sample.” [0013] “A probability density function of the GMM is based on a Gaussian distribution. An outlier cluster is defined based on a user-defined threshold. An outlier point is defined based on a user-defined confidence level.” Granson, [col. 15, lines 7-11] “The imputation phase addresses potential issues related to sensitivity problems of neural networks to missing values within portions of a data set, and also by addressing potential issues caused by statistically significant outlier values.” – Cao’s successive recurrent computations in the forward direction correspond to the claimed first plurality of recurrent neural network cells arranged in a forward layer. Cao generates estimated values, incorporates the estimated values into complement input, and uses the complement input to predict the recurrent state for the next recurrent computation. Thus, a predicted output value generated by a first forward recurrent cell is provided as input to a second forward recurrent cell. The corresponding predicted extremeness score is determined as mapped above through Jung and Granson and is incorporated with the predicted output value into a successive recurrent computation.); and a second plurality of recurrent neural network cells arranged in a backward layer, wherein a second recurrent neural network cell of the second plurality of recurrent neural network cells provides a predicted output value and a predicted extremeness score as an input to a first recurrent neural network cell of the second plurality of recurrent neural network cells (Cao, [section 4.2] “In this section, we propose an improved version called BRITS-I. The algorithm alleviates the above-mentioned issues by utilizing the bidirectional recurrent dynamics on the given time series, i.e., besides the forward direction, each value in time series can be also derived from the backward direction by another fixed arbitrary function…Consider the backward direction of the time series. In bidirectional recurrent dynamics, the estimation x ^ 4 reversely depends on x ^ 5 to x ^ 7 …Formally, the BRITS-I algorithm performs the RITS-I as shown in Eq. (1) to Eq. (5) in forward and backward directions, respectively…Similarly, in the backward direction, we obtain another estimation sequence { x ' ^ 1 ,   x ' ^ 2 , … , x ^ ' T } and another loss sequence { l 1 ' , l 2 ' … , l T ' } ” – Cao’s successive recurrent computations in the backwards direction correspond to the claimed second plurality of recurrent neural network cells arranged in a backward layer. Cao explains that an earlier version in the time series reversely depends on values occurring later in the time series and performs the same recurrent imputation operations in the backward direction. Thus, a predicted output value generated by a second backward recurrent cell is provided as input to a first backward recurrent cell. The corresponding predicted extremeness score is determined as mapped above through Jung and Granson and is incorporated with the predicted output value into the successive backward recurrent computation.), wherein a first predicted output value data and a first predicted extremeness score data of the forward layer is compared to a second predicted output value data and a second predicted extremeness score of the backward layer to identify the substitute value data (Cao, [section 4.2] “. In the forward direction, we obtain the estimation sequence { x ^ 1 ,   x ^ 2 , … , x ^ T } and the loss sequence { l 1 , l 2 , … , l T } . Similarly, in the backward direction, we obtain another estimation sequence { x ' ^ 1 ,   x ' ^ 2 , … , x ^ ' T } and another loss sequence { l 1 ' , l 2 ' … , l T ' } . We enforce the prediction in each step to be consistent in both directions by introducing the “consistency loss” … The final estimation in the t -th step is the mean of x ^ t , and x ' ^ t ” – Cao compares the predicted output values generated in the forward and backward layers using a consistency loss that evaluates the discrepancy between the bidirectional estimates. Cao then uses the mean of the forward and backward values to determine the final imputed value, corresponding to identifying the substitute-value data. In the combined system, the corresponding predicted extremeness scores mapped above through Jung and Granson are likewise compared for consistency between the forward and backward layers when identifying the substitute value data.); Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the teachings of Fouladgar, Granson, Jung, Chen, and Cao, wherein Fouladgar provides the time-series imputation framework using a recurrent neural network, Granson provides extremeness scoring and consideration of statistically significant outlier values during neural network imputation, Jung provides determining a normality threshold based on statistical measures including mean and standard deviation, Chen provides retail time-series data and retail forecasting, and Cao provides recurrent neural network cells operating in forward and backward layers, passing predicted values through successive recurrent computations, and comparing the forward and backward predictions to determine substitute value data. One would have been motivated to incorporate Granson’s extremeness scoring and Jung’s normality threshold determination into the recurrent imputation framework of Fouladgar, and to implement the framework using Cao’s bidirectional recurrent processing, so the predictions made from both temporal directions are evaluated for consistency and statistically extreme values are accounted for when determining substitute values. One would have further been motivated to apply the combined imputation technique to Chen’s retail time-series data to improve the accuracy and reliability of the reconstructed data used to train the model and generate the retail order volume forecast. Regarding claim 21, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view Cao teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao teaches: wherein the recurrent neural network cells of the first plurality of recurrent neural network cells and the second plurality of recurrent neural network cells each include a recurrent layer and a regression layer (Cao, [section 4.1.1] “We introduce a recurrent component and a regression component for imputation. The recurrent component is achieved by a recurrent neural network and the regression component is achieved by a fully-connected network… Eq. (1) is the regression component which transfers the hidden state h t - 1 to the estimated vector x ^ t . In Eq. (2), we replace missing values in x t with corresponding values in   t , and obtain the complement vector x t c …In Eq. (4), based on the decayed hidden state, we predict the next state   h t ” [section 4.2] “Formally, the BRITS-I algorithm performs the RITS-I as shown in Eq. (1) to Eq. (5) in forward and backward directions, respectively.” – Cao’s recurrent component corresponds to the claimed recurrent layer, and Cao’s fully connected regression component corresponds to the claimed regression layer. Cao performs the same recurrent and regression operations in both the forward and backward directions. Thus, the recurrent neural network cells of the first and second pluralities include a recurrent layer and a regression layer.). Claims 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Fouladgar et al., (NPL: “A Novel LSTM for Multivariate Time Series with Massive Missingness” (Published: 2020)) in view of Chen at al., (Pub. No.: US 20120303411 A1 (Filed: 2011)) further in view of Granson et al., (Pat. No.: US 11,816,539 B1 (Filed: 2017)) further in view of Jung (Pub. No.: US 20220164688 A1 (Filed: 2020) further in view of Zhiyong et al., (NPL: “Stacked bidirectional and unidirectional LSTM recurrent neural network for forecasting network-wide traffic state with missing values.” (Published: 2020)) further in view of Cao et al., (NPL: “BRITS: Bidirectional Recurrent Imputation for Time Series” (Published: 2018)). Regarding claim 4, Fouladgar in view Chen further in view of Granson further in view of Jung further in view of Cao teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. However, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao does not teach: wherein the recurrent neural network includes two separate bi-directional Long Short­ Term Memory network. However, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao further view of Zhiyong teaches the limitation: wherein the RNN includes two separate bi-directional Long Short­ Term Memory network. (Zhiyong, [Introduction] “The evaluation of prediction capability of stacked LSTM or BDLSTM – based models has great potential to facilitate the design of neural network models for traffic prediction.” Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Fouladgar, Chen, Granson, Jung, and Zhiyong before them, to incorporate two separate bi-directional Long Short-Term Memory network. One would have been motivated to use two separate bi-directional LSTM networks in order to model past and future speed values at multiple locations simultaneously, capturing both upstream and downstream traffic dependencies more accurately. Regarding claim 14, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. However, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao does not teach: wherein the RNN may include two separate bi-directional Long Short­ Term Memory network. However, Fouladgar in view of Chen further in view of Granson further in view of Jung further in view of Cao further view of Zhiyong teaches the limitation: wherein the RNN may include two separate bi-directional Long Short­ Term Memory network. (Zhiyong, [Introduction] “The evaluation of prediction capability of stacked LSTM or BDLSTM – based models has great potential to facilitate the design of neural network models for traffic prediction.” Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Fouladgar, Chen, Granson, Jung, and Zhiyong before them, to incorporate two separate bi-directional Long Short-Term Memory network. One would have been motivated to use two separate bi-directional LSTM networks in order to model past and future speed values at multiple locations simultaneously, capturing both upstream and downstream traffic dependencies more accurately. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to whose telephone number is (571)272-6324. The examiner can normally be reached Mon - Thurs 7 AM - 5 PM, Every other Friday 7 AM - 4PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li B Zhen can be reached at 571-272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Daravanh Phakousonh/Examiner, Art Unit 2121 /Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121
Read full office action

Prosecution Timeline

Show 7 earlier events
Mar 03, 2026
Request for Continued Examination
Mar 12, 2026
Response after Non-Final Action
Apr 07, 2026
Non-Final Rejection mailed — §103
May 22, 2026
Interview Requested
Jun 03, 2026
Applicant Interview (Telephonic)
Jun 03, 2026
Examiner Interview Summary
Jun 25, 2026
Response Filed
Sep 22, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12572821
ACCURACY PRIOR AND DIVERSITY PRIOR BASED FUTURE PREDICTION
4y 0m to grant Granted Mar 10, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
25%
Grant Probability
99%
With Interview (+100.0%)
3y 3m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 4 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month