Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
Claims 1, 5, 6, 10, 12, 18, 20 - 22, 24 are amended. Claims 1 – 25 are pending and examined herein.
Claims 24, 25 are rejected under 35 U.S.C. 112(b).
Claims 1 – 25 are rejected under 35 U.S.C. 103.
Response to Amendment
The amendment filed June 3rd, 2026 has been entered. Claims 1, 5, 6, 10, 12, 18, 20 - 22, 24 are amended. Claims 1 – 25 are pending and examined herein. Applicant’s amendments to the abstract of the disclosure have overcome the objection to specification previously set forth in the Non-Final Rejection Office Action mailed March 3rd, 2026.
Response to Arguments
Applicant’s arguments, see pages 10-15, filed June 3rd, 2026, with respect to 35 U.S.C. § 101 rejection have been fully considered and are persuasive. Although the claims recite time series data processing and prediction, the amended independent claims now require particular limitations describing how the input data are structured and processed by the transformer model, including patch processing, single channel input tokens, and masked patch pretraining. The specification paragraphs [20-21, 51] cited by applicant explains that these limitations reduce computational complexity, avoid noisy mixing across channels, and improve operation of the transformer based forecasting models. Considered as a whole, the amended limitations reflect the disclosed technological improvement and integrate the recited abstract idea into a practical application. Therefore, the 35 U.S.C. 101 rejection of claims 1 – 25 has been withdrawn.
Applicant’s arguments, see pages 16-18, filed June 3rd, 2026, with respect to 35 U.S.C. § 103 rejection have been fully considered and are persuasive. The cited references do not fairly teach or suggest the claim as amended. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Zhang et al. (NPL: “Crossformer: Transformer Utilizing Cross-Dimension Dependency for Multivariate Time Series Forecasting”, 2023) and Vaswani et al. (NPL: “Attention is all you need”, 2017). See rejection below.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 24, 25 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 24, 25 recites the limitation “applies scaled production”. It does not identify what quantity is scaled, what operation is “production”, or how the operation uses the previously recited matrices. Although the specification repeats the phrase “scaled production”, it does not define the phrase or otherwise establish its scope. Since the claims and specification do not presently recite or define the operation, the metes and bounds of “scaled production” is unclear. For examination purposes, “applies scaled production” in claims 24,25 are interpreted as “scaled dot product attention”, which is a known operation using previously recited query, key, and value matrices in transformer architecture.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 3 – 12, 14 – 23 are rejected under 35 U.S.C. 103 as being unpatentable over Tang et al. (NPL:” MTSMAE: Masked Autoencoders for Multivariate Time-Series Forecasting”) in view of Arik et al. (U.S. Pub. 2024/0249192 A1), Jawed et al. (NPL:” GQFormer: A Multi-Quantile Generative Transformer for Time Series Forecasting”), further in view of Zhang et al. (NPL: “CROSSFORMER: TRANSFORMER UTILIZING CROSS DIMENSION DEPENDENCY FOR MULTIVARIATE TIME SERIES FORECASTING” ).
Regarding Claim 1, Tang teaches
… wherein the dividing comprises aggregating time steps into patches. (Pg. 3 of Tang states “The original MAE continues the idea of ViT, processing image data X 2 RH_W_C, where (H;W) is the resolution of the original image, C is the number of channels. For MTSD, X 2 RLx_dx , the original patch embedding method is no longer applicable. Therefore, unlike the method of patch image data in ViT, we patch MTSD in the direction of time after embedding:
PNG
media_image1.png
25
278
media_image1.png
Greyscale
PNG
media_image2.png
21
281
media_image2.png
Greyscale
where the kernel width of one-dimensional convolutional filter and stride = P, Xpte is the final result of the patch embedding. We use two-step one-dimensional convolution to control the resolution of each MSTD patch to (P; dmodel), so the length of the last input sequence L = Lx=P2.” Tang teaches temporal patching where the input is segmented into patches via one dimensional convolution. )
transforming, by the processor set, the set of univariate time subseries into a univariate prediction result series using a transformer model; (Pg. 3 C. Encoder and Decoder section of Tang states “Our encoder is the encoder of transformer. In the pre-training, our encoder embeds only visible, unmasked patches through patch embedding, and then processes the output data through a series of transformer encoder blocks.” Pg. 2 III Methodology section of Tang states “The problem of multivariate time-series forecasting is to input the past sequence Xt= xt1,··· ,xtLx|xti∈Rdx at time t, and output the predict the corresponding future sequence Yt = yt1,··· ,ytLy|yti∈Rdy , where Lx and Ly are the lengths of input and output sequences respectively, and dx and dy are the feature dimensions of input X and output Y respectively. Our masked autoencoders (MTSMAE) is a simple autoencoding method and the training process is divided into two stages, as shown in the Fig. 1.” Since the multivariate forecast output in R3 is defined over feature dimensions, the predicted future values for any single feature constitute a univariate prediction result series. Thus, applying the transformer encoder/decoder to each univariate subseries yields the required univariate prediction series.)
and outputting, by the processor set, the multivariate predictive result for providing time series forecasting to a second system. (Pg. 2 III Methodology section of Tang states “The problem of multivariate time-series forecasting is to input the past sequence Xt= xt1,··· ,xtLx|xti∈Rdx at time t, and output the predict the corresponding future sequence Yt = yt1,··· ,ytLy|yti∈Rdy , where Lx and Ly are the lengths of input and output sequences respectively, and dx and dy are the feature dimensions of input X and output Y respectively. “ Tang outputs predicted future data and under BRI could be outputting the forecast result for downstream by another system.)
However, Tang does not explicitly teach
A method, comprising: receiving, by a processor set, an input time series from an external device in a first system;
dividing, by the processor set, the input time series to a set of univariate time subseries;
concatenating, by the processor set, the univariate prediction result series to a multivariate predictive result;
wherein each input token to the transformer model contains information from only a single channel of the input time series;
Arik teaches that
A method, comprising: receiving, by a processor set, an input time series from an external device in a first system; ([0020] of Arik states “The system can receive input data 105, which can be represented according to any of a variety of data structures, including vectors, tables, matrices, tensors, and so on. The input data 105 can include multiple data points, each point corresponding to a point in time or time step.” [0043] of Arik states “The system receives one or more input data points, according to block 405. Each input data point corresponds to a respective past time step earlier in time than a current time step. The current time step can vary depending on the input data and/or the time at which the system receives the input data.”)
Jawed teaches that
dividing, by the processor set, the input time series to a set of univariate time subseries; (Pg. 3 III Background section of Jawed states “We consider N related univariate time series data Y∈RT×N where each time series Yn∈RT is noted for a total of t=[1,..τ,...,T] timesteps. The variable τ is used to indicate The partitioning of the conditioning and the forecasting ranges.” Dividing the input timeseries Y into a set of univariate time subseries comprises extracting each channel/feature sequence Yn to obtain the set)
concatenating, by the processor set, the univariate prediction result series to a multivariate predictive result; (Pg. 3 III Background section of Jawed states “We consider N related univariate time series data Y∈RT×N where each time series Yn∈RT is noted for a total of t=[1,..τ,...,T] timesteps. The variable τ is used to indicate the partitioning of the conditioning and the forecasting ranges. In addition to the real-world time series we also consider C many social time covariates X ∈ RT×C that are observed in the entire range. We aim to model the following conditional distribution: p(Y n τ+1:T |Y n 1:τ ,X1:T , Θ) (1) This formulation in Eq. 1 explicitly models for multiple tasks jointly conditioned on the same input and model parameters Θ. This is in contrast to other works that reduce the problem complexity by formulating a simpler single step forecasting task p(Y nτ+1|Y n 1:τ ,X1:τ+1, Θ)4. Note that our formulation and following background is similar to [23].” Jawed represents the multivariate time series as N univariate series assembled as Y∈RT×N and further predicts a future trajectory per univariate series in Equation 1 Yn τ+1under a formulation that models for multiple tasks jointly. Accordingly, the set of per series predicted trajectories is assembled (i.e. concatenated across the N series dimension) to form the multivariate predictive result.)
Zhang teaches that
wherein each input token to the transformer model contains information from only a single channel of the input time series; (To this end, we propose Dimension-Segment-Wise (DSW) embedding where the points in each dimension are divided into segments of length Lseg and then embedded:
PNG
media_image3.png
72
283
media_image3.png
Greyscale
where x(s)i,d ∈ RLseg is the i-th segment in dimension d with length Lseg. For convenience, we assume that T,τ are divisible by Lseg3. Then each segment is embedded into a vector using linear projection added with a position embedding:
PNG
media_image4.png
21
142
media_image4.png
Greyscale
where E ∈ Rdmodel_Lseg denotes the learnable projection matrix, and E(pos)i;d 2 Rdmodel denotes the learnable position embedding for position (i; d). After embedding, we obtain a 2D vector array H = {h_i,d|1<= i<= T/L_seg, 1 <=d <= D}, where each h_i,d represents a univariate time series segment. The idea of segmentation is also used in Du et al. (2022), which splits the embedded 1D vector sequence into segments to compute the Segment-Correlation in order to enhance locality and reduce computation complexity.” Zhang teaches DSW embedding where the point in each dimension are divided into temporal segments and each embedded vector represents a univariate time series segment drawn from only one dimension D, which corresponds to one channel)
It would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to combine the teachings of Tang, Jawed, Arik, Zhang. Tang teaches transformer based time series modeling using patch embedding, including masking a subset of patches, pre-training by reconstructing masked patches, and then using the transformer in fine-tuning to output prediction series. Arik teaches a forecasting workflow that receives past time step inputs and generates future time step predicted outputs, and further teaches weight sharing across features, supporting efficient forecasting when multiple channels/features are present. Jawed teaches a forecasting architecture for time series that includes applying a flatten operation to learned embeddings and using a shared linear head to produce forecast outputs. Zhang teaches dimension segment wise embedding where the points in each dimension are divided into temporal segments and each embedded vector represents a univariate time series segment using only one dimension D. One with ordinary skill in the art would be motivated to incorporate the teachings of Zhang, Jawed and Arik with Tang because these are structurally compatible design choice in transformer forecasting systems by enabling efficient pretraining, straightforward and widely used linear design, and ensuring consistent, scalable forecasting for multiple channels using past time series inputs. Zhang further preserves the time and dimension information of the input in addition so the resulting token is formed from neighboring timeseries values of only one dimension or channel. It would have been predictable combination to improve implementation and robustness of the model practically while restricting to using single channel input tokens.
Regarding Claim 3, the rejection of claim 1 is incorporated herein. Furthermore, the combination of Tang, Arik, Jawed, Zhang teaches
wherein the input time series comprises a multivariate time series. (Pg. 2 III Methodology section of Tang states “The problem of multivariate time-series forecasting is to input the past sequence Xt= xt1,··· ,xtLx|xti∈Rdx at time t, and output the predict the corresponding future sequence Yt = yt1,··· ,ytLy|yti∈Rdy , where Lx and Ly are the lengths of input and output sequences respectively, and dx and dy are the feature dimensions of input X and output Y respectively. “ Input timeseries comprise multivariate timeseries while doing multivariate timeseries forecasting. )
Regarding Claim 4, the rejection of claim 3 is incorporated herein. Furthermore, the combination of Tang, Arik, Jawed, Zhang teaches
wherein the multivariate time series comprises a multi-channel signal. (Pg. 2 III Methodology section of Tang states “The problem of multivariate time-series forecasting is to input the past sequence Xt= xt1,··· ,xtLx|xti∈Rdx at time t, and output the predict the corresponding future sequence Yt = yt1,··· ,ytLy|yti∈Rdy , where Lx and Ly are the lengths of input and output sequences respectively, and dx and dy are the feature dimensions of input X and output Y respectively. “ Multivariate with feature dimensions implies multiple channels/features, which is known standard interpretation.)
Regarding Claim 5, the rejection of claim 1 is incorporated herein. Furthermore, the combination of Tang, Arik, Jawed, Zhang teaches
wherein the transforming the set of univariate time subseries into the univariate prediction result series comprises normalizing and segmenting the univariate time subseries into the patches. ([0033] of Arik states “The system 100 receives input to the time mixing layer 302 and normalizes the input at a two-dimensional normalization layer (2D Norm) layer 310. At the 2D Norm layer 310, the system 100 normalizes over both time and feature dimensions of the input, to maintain a consistent scale between the time-mixing and feature-mixing operations at the later stages of the time mixing layer 302 and feature mixing layer 304, respectively.” Pg. 3 A. Patch embedding section of Tang states “The original MAE continues the idea of ViT, processing image data X∈ 2 RH*W*C, where (H;W) is the resolution of the original image, C is the number of channels. For MTSD, X ∈ RLx*dx , the original patch embedding method is no longer applicable. Therefore, unlike the method of patch image data in ViT, we patch MTSD in the direction of time after embedding: Xh = Conv1d(X) ∈ RLx=Pdmodel (6) Xpte = Conv1d(Xh) ∈ RLx=P2dmodel (7) where the kernel width of one-dimensional convolutional filter and stride = P, Xpte is the final result of the patch embedding.” Arik teaches normalizing the input and Tang teaches segmenting the time series into patch tokens.)
Regarding Claim 6, the rejection of claim 5 is incorporated herein. Furthermore, the combination of Tang, Arik, Jawed, Zhang teaches
wherein the patches are local and semantic information in the aggregated time steps. (Pg. 4 C. Encoder and Decoder section of Tang states “In the pre-training, our MTSMAE reconstructs the input by recovering the specific value of each masking patch. Each element output by the decoder is a vector that can represent a patch. The last layer of the decoder is a linear projection, whose output channel is P D, P is the length of the patch, and D is the dimension of the time-series. In the fine-tuning, each element output by the decoder represents the data yti , and the output channel of the last layer of linear projection is D.” Patches have a defined patch length P where each patch aggregates P timesteps. These patch tokens further serve as the learned representation units (i.e. semantic feature) for the transformer encoder/decoder.)
Regarding Claim 7, the rejection of claim 5 is incorporated herein. Furthermore, the combination of Tang, Arik, Jawed, Zhang teaches
wherein the transforming the set of univariate time subseries into the univariate prediction result series further comprises transforming the patches into a representation. (Pg. 2 III Methodology section of Tang states “Our masked autoencoders (MTSMAE) is a simple autoencoding method and the training process is divided into two stages, as shown in the Fig. 1. As all autoencoders, there is an encoder and a decoder in our method. The encoder maps the observed signal to a latent representation and the decoder reconstructs the original signal from the latent representation in the pre-training, or output Y in the fine-tuning.”)
Regarding Claim 8, the rejection of claim 7 is incorporated herein. Furthermore, the combination of Tang, Arik, Jawed, Zhang teaches
wherein the transforming the set of univariate time series into the univariate prediction result series further comprises utilizing a flatten layer with a linear head on the representation to obtain the univariate prediction result series. (Pg. 5 D. Decoder section of Jawed states “ξFlat˜y = Flatten(ξ˜y) (14)ξ1:M˜y = Repeat(ξFlat˜y ,M) (15)qαi,τ : = [ξi ˜y ξαi ]WMTL + bMTL ∀i = [1, ...,M] (16) In the above equations, we first flatten the embedding of the time series to one feature axis, this results into the embedding size: (dmodel×len(1 : τ )), where dmodel indicates the embedded dimensionality of each timestep input. Next we repeat these M many times to combine these with the quantile embeddings in Eq. 10. Observe that each of the [1, ...M] quantile embedding is different, but the time series embedding ξFlat˜y remains the same. Finally, a shared fully connected layer, given by parameters WMTL ∈R(dmodel×len(1:τ)+dmodel)×len(τ+1:T),bMTL ∈ Rlen(τ+1:T) is learned to produce a quantile forecast based on the concatenated repeated representation of the time series and the embeddings of the implicit quantile levels.” Applies flatten to the learned representation and uses linear layer as the prediction head to output forecasts.)
Regarding Claim 9, the rejection of claim 1 is incorporated herein. Furthermore, the combination of Tang, Arik, Jawed, Zhang teaches
wherein the univariate time subseries comprises a plurality of channel independent signals. (Pg. 3 III Background section of Jawed states “We consider N related univariate time series data Y∈RT×N where each time series Yn∈RT is noted for a total of t=[1,..τ,...,T] timesteps. The variable τ is used to indicate The partitioning of the conditioning and the forecasting ranges.”Jawed models N separate univariate series (i.e. channels) inside a multivariate structure.)
Regarding Claim 10, the rejection of claim 9 is incorporated herein. Furthermore, the combination of Tang, Arik, Jawed, Zhang teaches
wherein each of the plurality of channel independent signals have a same model weight as a weight of remaining channel independent signals. ([0028] of Arik states “In this specification, a “mixing layer” or “mixer layer” can refer to a layer of both time-domain and feature-domain operations. Additionally, a “time mixing layer” or “time mixer layer” can refer to a layer of time-domain operations, while a “feature mixing” or “feature mixer” layer can refer to a layer of feature-domain operations. Layers are collections of operations that at least partially depend on trainable weights or parameter values. Machine learning models may include different layers, such as fully-connected layers, dropout layers, etc.” [0006] of Arik states “The time series mixer includes MLPs that alternate between time-domain input and feature-domain input… Time-domain MLPs are reused or shared across all the features of an input time series, while feature-domain MLPs are reused or shared across all time steps of the input time series.” [0036] of Arik states “The time-mixing MLP 320 is shared across each feature of the transposed data 319. In other words, the system 100 processes values for each feature of the transposed data 319 through the time-mixing MLP 320.” Shared model components across all features mean same model weights are applied across channels/features.)
Regarding Claim 11, the rejection of claim 1 is incorporated herein. Furthermore, the combination of Tang, Arik, Jawed, Zhang teaches
wherein the transformer model comprises a supervised model. (Pg. 1 Abstract of Tang states “In this paper, according to the data characteristics of multivariate time-series, a patch embedding method is proposed, and we present an self-supervised pre-training approach based on Masked Autoencoders (MAE), called MTSMAE, which can improve the performance significantly over supervised learning without pre-training.”)
Regarding Claim 12, the combination of Tang, Arik, Jawed, and Zhang teaches
pre-train a transformer model using historically reconstructed masked patches (Pg. 4 C. Encoder and Decoder section of Tang states “In the pre-training, we set up the decoder as MAE. The input to the decoder is a complete set including the visible patches output by encoder and mask tokens, where as the vector of learning, masked tokens are the data to be recovered… In the pre-training, our MTSMAE reconstructs the input by recovering the specific value of each masking patch.” Tang pre-trains via masked patch reconstruction and performs forecasting via transformer encoder and prediction decoder.)
The rest of claim 12 recites substantially similar subject matter to claim 1 respectively and is rejected with the same rationale, mutatis mutandis.
Regarding claim 14, the rejection of claim 12 is incorporated herein. Regarding claim 15, the rejection of claim 14 is incorporated herein. Claims 14 – 15 recite substantially similar subject matter as claims 3 – 4 respectively, and are rejected with the same rationale, mutatis mutandis.
Regarding claim 16, the rejection of claim 15 is incorporated herein. Furthermore, the combination of Tang, Arik, Jawed, Zhang teaches
masking the univariate time subseries into masked patches and non-masked patches. (Pg. 3-4 C. Encoder and Decoder section of Tang states “Our encoder is the encoder of transformer. In the pretraining, our encoder embeds only visible, unmasked patches through patch embedding, and then processes the output data through a series of transformer encoder blocks. Our encoder only operates on a small part of the whole set, e.g., only 15%, which can greatly reduce the redundancy of information and increase the overall understanding of the model beyond low-level information. In the fine-tuning, our encoder can see all the patches… In the pre-training, we set up the decoder as MAE. The input to the decoder is a complete set including the visible patches output by encoder and mask tokens, where as the vector of learning, masked tokens are the data to be recovered.” Tang explicitly distinguishes unmasked patches vs masked patches (random masking) and reconstructs the masked ones in pretraining.)
The rest of claim 16 recites substantially similar subject matter to claim 5 respectively and is rejected with the same rationale, mutatis mutandis.
Regarding claim 17, the rejection of claim 16 is incorporated herein. Furthermore, the combination of Tang, Arik, Jawed, Zhang teaches
wherein the transforming the set of univariate time subseries into the univariate prediction result series further comprises utilizing a linear layer on the non-masked patches to obtain the univariate prediction result series. (Pg. 3-4 C. Encoder and Decoder section of Tang states “Our encoder is the encoder of transformer. In the pretraining, our encoder embeds only visible, unmasked patches through patch embedding, and then processes the output data through a series of transformer encoder blocks… In the pre-training, our MTSMAE reconstructs the input by recovering the specific value of each masking patch. Each element output by the decoder is a vector that can represent a patch. The last layer of the decoder is a linear projection, whose output channel is P D, P is the length of the patch, and D is the dimension of the time-series. In the fine-tuning, each element output by the decoder represents the data yti , and the output channel of the last layer of linear projection is D. Our loss function is calculated by the mean square error (MSE) between the model output data yo (recovery, prediction) and the real data y.” Tang feeds unmasked patches through the model and the decoder ends in a linear projection producing the prediction outputs.)
Regarding claim 18, the rejection of claim 16 is incorporated herein. Furthermore, the combination of Tang, Arik, Jawed, Zhang teaches
wherein the transforming the set of univariate time subseries into the univariate prediction results series further comprises reconstructing the masked patches. (Pg. 4 C. Encoder and Decoder section of Tang states “In the pre-training, our MTSMAE reconstructs the input by recovering the specific value of each masking patch.”)
Regarding claim 19, the rejection of claim 12 is incorporated herein. Regarding claim 20, the rejection of claim 19 is incorporated herein. Claims 19 – 21 recite substantially similar subject matter as claims 9 – 11 respectively, and are rejected with the same rationale, mutatis mutandis.
Regarding claim 23, the rejection of claim 22 is incorporated herein. Claims 22 – 23 recite substantially similar subject matter as claims 12 and 16 respectively, and are rejected with the same rationale, mutatis mutandis.
Claims 24 – 25 are rejected under 35 U.S.C. 103 as being unpatentable over Tang et al. (NPL:” MTSMAE: Masked Autoencoders for Multivariate Time-Series Forecasting”) in view of Arik et al. (U.S. Pub. 2024/0249192 A1), Jawed et al. (NPL:” GQFormer: A Multi-Quantile Generative Transformer for Time Series Forecasting”), further in view of Vaswani et al. (NPL: “Attention Is All You Need”).
Regarding Claim 24, the combination of Tang, Arik, Jawed teaches
receiving, by a processor set, a univariate time series; (Pg. 3 III Background section of Jawed states “We consider N related univariate time series data Y∈RT×N where each time series Yn∈RT is noted for a total of t=[1,..τ,...,T] timesteps. The variable τ is used to indicate The partitioning of the conditioning and the forecasting ranges.” [0020] of Arik states “The system can receive input data 105, which can be represented according to any of a variety of data structures, including vectors, tables, matrices, tensors, and so on. The input data 105 can include multiple data points, each point corresponding to a point in time or time step.” [0043] of Arik states “The system receives one or more input data points, according to block 405. Each input data point corresponds to a respective past time step earlier in time than a current time step. The current time step can vary depending on the input data and/or the time at which the system receives the input data.” Receiving timeseries comprised of univariate time series data.)
dividing, by the processor set, the univariate time series into a non-overlapped set of patches; (Pg. 3 A. Patch embedding section of Tang states “The original MAE continues the idea of ViT, processing image data X∈ 2 RH*W*C, where (H;W) is the resolution of the original image, C is the number of channels. For MTSD, X ∈ RLx*dx , the original patch embedding method is no longer applicable. Therefore, unlike the method of patch image data in ViT, we patch MTSD in the direction of time after embedding: Xh = Conv1d(X) ∈ RLx=Pdmodel (6) Xpte = Conv1d(Xh) ∈ RLx=P2dmodel (7) where the kernel width of one-dimensional convolutional filter and stride = P, Xpte is the final result of the patch embedding.” Dividing the series into time direction patches, also kernel width of one dimensional convolutional filter and stride = P means each patch window is exactly one patch step, which is non-overlapping.)
masking a subset of the non-overlapped set of patches to a masked patch series; (Pg. 3 of Tang states “A random masking method is adopt, that is, the patches are randomly sampled without replacement, and follow the uniform distribution. The random sampling can tremendously remove the information redundancy of MSTD by deleting a large number of patches (i.e., high masking rate)… In the pretraining, our encoder embeds only visible, unmasked patches through patch embedding, and then processes the output data through a series of transformer encoder blocks.”)
pre-training a transformer model using historically reconstructed masked patches; (Pg. 4 of Tang states “In the pre-training, our MTSMAE reconstructs the input by recovering the specific value of each masking patch. Each element output by the decoder is a vector that can represent a patch. The last layer of the decoder is a linear projection, whose output channel is P * D, P is the length of the patch, and D is the dimension of the time-series.”)
transforming, by the processor set, the non-overlapped set of patches into a representation using the pre-trained transformer model, (Pg. 2 III Methodology section of Tang states “Our masked autoencoders (MTSMAE) is a simple autoencoding method and the training process is divided into two stages, as shown in the Fig. 1. As all autoencoders, there is an encoder and a decoder in our method. The encoder maps the observed signal to a latent representation and the decoder reconstructs the original signal from the latent representation in the pre-training, or output Y in the fine-tuning.” Pg. 3 C. Encoder and Decoder section of Tang states “Our encoder is the encoder of transformer. In the pretraining, our encoder embeds only visible, unmasked patches through patch embedding, and then processes the output data through a series of transformer encoder blocks”)
wherein the pre-trained transformer model generates the representation by transforming the non-overlapped set of patches… applies scaled production and a feed forward network with residual connections. (Pg. 4 of Tang states “In the fine-tuning, our encoder can see all the patches. Except for this, it is no different from the encoder in the pre-training. The Transformer encoder layers are composed of two sub-blocks. The first is a multi-head self-attention mechanism (MSA), and the second is a simple, position-wise fully connected feed-forward network (MLP). Residual connections [30] are used around each of the two subblocks , and layer normalization (LN) [31] is then performed.”)
obtaining, by the processor set, a univariate prediction result series by using a flatten layer with a linear head on the representation; (Pg. 5 D. Decoder section of Jawed states “ξFlat˜y = Flatten(ξ˜y) (14)ξ1:M˜y = Repeat(ξFlat˜y ,M) (15)qαi,τ : = [ξi ˜y ξαi ]WMTL + bMTL ∀i = [1, ...,M] (16) In the above equations, we first flatten the embedding of the time series to one feature axis, this results into the embedding size: (dmodel×len(1 : τ )), where dmodel indicates the embedded dimensionality of each timestep input. Next we repeat these M many times to combine these with the quantile embeddings in Eq. 10. Observe that each of the [1, ...M] quantile embedding is different, but the time series embedding ξFlat˜y remains the same. Finally, a shared fully connected layer, given by parameters WMTL ∈R(dmodel×len(1:τ)+dmodel)×len(τ+1:T),bMTL ∈ Rlen(τ+1:T) is learned to produce a quantile forecast based on the concatenated repeated representation of the time series and the embeddings of the implicit quantile levels.” Applies flatten to the learned representation and uses linear layer as the prediction head to output forecasts.)
and outputting, by the processor set, the univariate prediction result series. (Pg. 2 III Methodology section of Tang states “The problem of multivariate time-series forecasting is to input the past sequence Xt= xt1,··· ,xtLx|xti∈Rdx at time t, and output the predict the corresponding future sequence Yt = yt1,··· ,ytLy|yti∈Rdy , where Lx and Ly are the lengths of input and output sequences respectively, and dx and dy are the feature dimensions of input X and output Y respectively. “ [0044] of Arik states “In processing the one or more input data points as described herein, the system generates one or more output data points. Each output data point corresponds to a respective future time step later in time than the current time step, and each output data point including respective predicted values for one or more of the features at the respective future time step.”)
However, the combination does not explicitly teach that
transforming… Query matrics, key matrics, value matrices, and applies scaled production and a feed forward network...
Vaswani teaches that
transforming… Query matrics, key matrics, value matrices, and applies scaled production and a feed forward network with residual connections. (Pg. 2 – 3 of Vaswani states “The first is a multi-head self-attention mechanism, and the second is a simple, position- wise fully connected feed-forward network. We employ a residual connection [10] around each of the two sub-layers, followed by layer normalization [1]… An attention function can be described as mapping a query and a set of key-value pairs to an output, where the query, keys, values, and output are all vectors.” Pg. 3 – 4 of Vaswani states “We call our particular attention "Scaled Dot-Product Attention" (Figure 2). The input consists of queries and keys of dimension dk, and values of dimension dv… In practice, we compute the attention function on a set of queries simultaneously, packed together into a matrix Q. The keys and values are also packed together into matrices K and V . We compute the matrix of outputs as:
PNG
media_image5.png
51
272
media_image5.png
Greyscale
” )
It would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to combine the teachings of Vaswani with Tang, Jawed, Arik. Tang teaches transformer based time series modeling using patch embedding, including masking a subset of patches, pre-training by reconstructing masked patches, and then using the transformer in fine-tuning to output prediction series. Arik teaches a forecasting workflow that receives past time step inputs and generates future time step predicted outputs, and further teaches weight sharing across features, supporting efficient forecasting when multiple channels/features are present. Jawed teaches a forecasting architecture for time series that includes applying a flatten operation to learned embeddings and using a shared linear head to produce forecast outputs. Vaswani teaches generating query, key, and value matrices by learned linear projection and computing attention as a scaled dot product of these matrices. One with ordinary skill in the art would be motivated to incorporate the teachings of Vaswani with Jawed, Arik, and Tang because that formulation is a well-known implementation of transformer attention. It would have been predictable combination to merely apply known method to the attention operation already taught in Tang.
Claim 25 recite substantially similar subject matter as claim 24 respectively, and is rejected with the same rationale, mutatis mutandis.
Claims 2, 13 are rejected under 35 U.S.C. 103 as being unpatentable over Tang et al. (NPL:” MTSMAE: Masked Autoencoders for Multivariate Time-Series Forecasting”) in view of Arik et al. (U.S. Pub. 2024/0249192 A1), Jawed et al. (NPL:” GQFormer: A Multi-Quantile Generative Transformer for Time Series Forecasting”), Zhang et al. (NPL: “CROSSFORMER: TRANSFORMER UTILIZING CROSS DIMENSION DEPENDENCY FOR MULTIVARIATE TIME SERIES FORECASTING” ), further in view of Shabani et al. (U.S. Pub. 2023/0368002 A1).
Regarding claim 2, the rejection of claim 1 is incorporated herein. Furthermore, the combination of Tang, Arik, Jawed, and Zhang does not explicitly teach
wherein the external device comprises a smart sensor, the first system comprises a manufacturing system, and the second system comprises a planning system in communication with the first system.
However, Shabani teaches that
wherein the external device comprises a smart sensor, the first system comprises a manufacturing system, and the second system comprises a planning system in communication with the first system. ([0002] of Shabani states “Time Series Forecasting is among the most well-known problems in many domains such as sensor network monitoring, traffic and economics planning, astronomy, economic and financial forecasting, inventory planning, and weather and disease propagation forecasting.” It is conventional predictable application of the same forecasting pipeline.)
It would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to combine the teachings of Shabani with the combination of Tang, Jawed, Arik, Zhang. Tang teaches transformer based time series modeling using patch embedding, including masking a subset of patches, pre-training by reconstructing masked patches, and then using the transformer in fine-tuning to output prediction series. Arik teaches a forecasting workflow that receives past time step inputs and generates future time step predicted outputs, and further teaches weight sharing across features, supporting efficient forecasting when multiple channels/features are present. Jawed teaches a forecasting architecture for time series that includes applying a flatten operation to learned embeddings and using a shared linear head to produce forecast outputs. Zhang teaches dimension segment wise embedding where the points in each dimension are divided into temporal segments and each embedded vector represents a univariate time series segment using only one dimension D. Shabani teaches complementary timeseries forecasting processing, such as handling multiple component series, to reinforce the forecasting implementation and provide various use of time series forecasting in different systems. One with ordinary skill in the art would be motivated to incorporate the teachings of Shabani with the combination of Tang, Jawed, Arik, Zhang because it is adding known time series processing techniques to improve robustness and applicability across different time series applications. It would have been predictable combination to merely apply the combined system for compatibility across different platforms.
Regarding claim 13, the rejection of claim 12 is incorporated herein. Claim 13 recites substantially similar subject matter as claim 2 respectively, and is rejected with the same rationale, mutatis mutandis.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BYUNGKWON HAN whose telephone number is (571)272-5294. The examiner can normally be reached M-F: 9:00AM-6PM PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li B Zhen can be reached at (571)272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BYUNGKWON HAN/ Examiner, Art Unit 2121
/Li B. Zhen/ Supervisory Patent Examiner, Art Unit 2121