DETAILED ACTION
1. This office action is in response to the Application No. 18918077 filed
on 05/21/2026. Claims 4, 6, 9, 10 and 12 has been cancelled. Claims 1-3, 5, 7, 8 and 11 are presented for examination and are currently pending.
Priority
2. The Examiner notes that the following applications 18737906 filed 06/07/2024, 18736498 filed 06/06/2024, 63651359 filed 05/23/2024 has no support for the following limitations:
-“emerging patterns not represented by codewords”,
-“monitoring usage frequencies of the codewords”
-“generating new codewords for the identified emerging patterns”
-“combining the extracted features and interactions through a fusion layer to produce a fused vector”
-“a short-term forecast for the numerical time series data”.
Furthermore, the above listed applications appears to disclose “frequency”, but
it does not appear to disclose “monitoring usage frequencies of the codewords”.
In addition, the above listed applications appears to disclose “codewords” and “patterns” but it does not appear to disclose “monitoring usage frequencies of the codewords” or “generating new codewords for the identified emerging patterns”.
Application 18737906 filed 06/07/2024 has support for “combining a portion of the original time series data points with a set of truncated data points and a sequence of zeros”. It does not appear to disclose “combining the extracted features and interactions through a fusion layer to produce a fused vector”.
Application 18736498 filed 06/06/2024 has support for “combined input sequence contains information from both the numerical and text input data”. It does not appear to disclose “combining the extracted features and interactions through a fusion layer to produce a fused vector”.
Application 63651359 filed 05/23/2024 has support for “The combined input sequence contains information from both the numerical and text input data”, “learn patterns within the codeword sequences”, “assigns codewords based on the frequency of occurrence of each symbol”. It does not appear to disclose “emerging patterns not represented by codewords”, or “monitoring usage frequencies of the codewords”, or “combining the extracted features and interactions through a fusion layer to produce a fused vector” or “a short-term forecast for the numerical time series data”.
As a result, for the purpose of prosecution, the effective filling date for application
18918077 filed 10/17/2024 has been used for the prior art rejection.
Response to Arguments
3. The amended claimed filed 05/21/2026 has overcome the 101 rejection. Furthermore, the Applicant’s argument on page 10 that “The technical problem is representation drift in codebook-based systems processing evolving multi- modal data streams (see [0379]-[0384] of the specification)”. According to [0379] of the instant specification: “an adaptive codebook generation system improves the model’s ability to maintain relevance and accuracy in the fast-paced and ever-evolving financial markets. The system receives new market data 2700, which could encompass a wide range of financial information including real-time trading data, breaking news, economic indicators, and social media sentiment related to financial markets. This continuous stream of data is essential for keeping the model attuned to the latest market trends and events”.
These arguments are persuasive because it improves the technological field of adaptive systems in financial market. As a result, the 101 rejection is withdrawn.
The amended claimed filed 05/21/2026 has overcome the 112(b) rejection and as a result, the 112(b) has been withdrawn.
Applicant’s arguments regarding the prior art rejection are moot in view of the new grounds of rejection. The Examiner is withdrawing the rejections in the previous Office Action because the Applicant’s amendments necessitated the new grounds of rejection presented in this Office Action.
However, Harikumar which was used in the previously in the art rejection is still relevant and applied.
Harikumar still teaches a deep learning system (In some implementations, a CNN is used to predict image tokens [0033]) for real-time time series forecasting (predict the next image tokens autoregressively (one token at a time in serial fashion [0041]) using a compound large codeword model (As used herein, in the context of this disclosure, a “codebook” is a visual dictionary of a set of visual words (or codewords) that represent one or more feature vectors of one or more images [0035])
allocate codewords to each data input, wherein codewords are mapped to a corresponding codebook (In some embodiments, a set of image tokens each having a unique integer value is obtained based on the codebook 502 a for the image encoder 510 and the codebook 502 b for the sketch encoder 520 [0070], Fig. 5. The Examiner notes a codebook contains codeword), and
wherein the codewords and their corresponding codebooks are adaptively updated to reflect incoming data inputs (codebook 502a is adaptively updated by incoming data inputs from 504a-n and codebook 502b is adaptively updated by incoming data inputs from 504b, Fig.5)
a machine learning core comprising a transformer-based architecture (Transformer models 106 and 604 are examples of the transformer model [0111]), then,
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lai to incorporating the teachings of Harikumar for the benefits of using a transformer model which increases the resolution of the reconstructed data that results in the intricate patterns being created without loss of fidelity and at high efficiency (Harikumar [0099])
The dependent claims 2, 3, 5, 8 and 11 which depend directly or indirectly from independent claims 1 and 7 are not patentable because the instant claims are still obvious over the prior art of record.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
4. Claims 1-3, 7 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. ("ReCTSi: resource-efficient correlated time series imputation via decoupled pattern learning and completeness-aware attentions." Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2024) in view of Harikumar et al. (US20230419551)
Regarding claim 1, Lai teaches a deep learning system (Deep learning has enabled sophisticated models that improve CTS imputation by capturing temporal and spatial patterns, abstract) for real-time time series forecasting (One school of deep models is based on the common framework used for CTS (Correlated Time Series) forecasting, pg. 1476, left col., third para.) using a compound large codeword model (To achieve this, we design a novel Multi-view Learnable Codebook (MvLC) mechanism, pg. 1477, left col., fourth para.),
receive a plurality of data inputs (CTS Data Input, Fig. 1, pg. 1474) comprising at least textual data (Monday... Fig. 2, pg. 1475) and numerical time series data (12:00, 13:00 ..., Fig. 2: Multi-view patterns in CTS (pg. 1475); spatio-temporal and associate each with persistent and transient components, as shown in Figure 2. Temporal patterns encompass time-related features; persistent ones recur periodically (e.g., weekly traffic flow variations), pg. 1475, right col., first para.);
allocate codewords to each data input (PT-pattern Codebook (CBPT): Persistent Temporal Patterns (PT patterns) represent periodic information that can employ fixed pattern representation rules for all temporal segments, ... We assign aPT-pattern tensor vPT𝑡 to each timestamp 𝑡 and broadcast it across different time series, pg. 1477, right col., third para.),
wherein codewords are mapped to a corresponding codebook (MvLC Mechanism. The Learnable Codebook serves as the foundation for MvLC. It functions as a dictionary that maps input information to fixed-length pattern tensors, pg. 1477, right col., first full para.), and
wherein the codewords and their corresponding codebooks (PT-pattern Codebook, PS-pattern Codebook, PST-pattern Codebook, pg. 1477, right col.) are adaptively updated to reflect incoming data inputs by (To achieve this, we design a novel Multi-view Learnable Codebook (MvLC) mechanism, ... This mechanism learns and stores representations of three subtle views of persistent patterns during training, allowing for easy retrieval during imputation via periodic information updates, pg. 1477, left col. fourth para.);
analyzing the incoming data inputs to identify emerging patterns not represented by codewords of the corresponding codebook (An optimal window size of T = 24 hours aligns with the dataset’s daily periodicity, thereby optimizing the model’s capabilities, pg. 1482, left col., first para. The Examiner notes daily periodicity indicates daily recurring incoming data at regular, predictable intervals);
monitoring usage frequencies of the codewords of the corresponding codebook (After training, the codebooks of ppe are stored and used for subsequent imputation inference via table look-ups, pg. 1479, right col., third to the last para. The Examiner notes that using codebooks for subsequent imputation inference indicates the usage of frequency of the codebooks are monitored);
generating new codewords for the identified emerging patterns (Similar to the PT-pattern codebook, the Persistent Spatial Pattern (PS-pattern) codebook aims to capture patterns that remain constant over time but vary spatially. This could be geographical location or identity information of timeseries... These values are then broadcast across timestamps to achieve the PS-pattern, pg. 1477, right col., second to the last para. The Examiner notes new codewords are generated when codebook captures patterns that remain constant over time but vary spatially); and
pruning, from the corresponding codebook (Initially, we evaluate each codebook individually by systematically removing them, pg. 1480, right col., second to the last para.),
codewords whose monitored usage frequencies fall below a threshold (spatio-temporal attention requires 4096 times more computational and memory resources than does spatial attention alone, which is prohibitive. To achieve a resource-efficient model, we therefore limit the consideration of ST-patterns to the ppe phase, pg. 1479, left col., 4.3.2 Completeness-aware Attention);
fuse codewords of the textual data and codewords of the numerical time series data (PT-pattern Codebook, PS-pattern Codebook, PST-pattern Codebook are fused together (Fig. 3, pg. 1478); Monday...(day of the week); 12:00, 13:00...(time of day), geographical location (spatial pattern), Fig. 2: Multi-view patterns in CTS, pg. 1475), which are dissimilar data types (PT-pattern Codebook (CBPT), Persistent Temporal-pattern (PT-pattern)... For instance, information such as day-of-the-week (DoW) or time-of-the-day (ToD) is inherently periodic, pg. 1477, right col., third to last para.; PS-pattern Codebook ... aims to capture patterns that remain constant over time but vary spatially. This could be geographical location, pg. 1477, right col., third to last para.), together into a single codeword representation (To emphasize more significant persistent patterns, we concatenate all three views of persistent patterns along the feature embedding dimension, pg. 1478, left column, second para.) by extracting features from the codewords of each data type (This phase utilizes primarily the MvLC to extract persistent pat terns from three distinct views: the Persistent Temporal Pattern (PT-pattern), the Persistent Spatial Pattern (PS-pattern), and the Persistent Spatio-temporal Pattern (PST-pattern) through learnable codebooks, pg. 1477, left col., last para.), capturing interactions between the extracted features (pattern learning architecture that represents are thinking of the pattern extraction process. CTS typically encompasses two types of patterns: persistent patterns, capturing how data points evolve in relation to static information, such as temporal cycles and the geographic location of time series, and transient patterns, describing fluctuations at specific timestamps and semantic spatial correlations, pg. 1477. Left col., second full para.), and
combining the extracted features and interactions through a fusion layer to produce a fused vector (Pattern Compression and Fusion Layer in Fig. 3 produces a fused vector);
process the fused vector representing the single codeword representation through a machine learning core (ReCTSi integrates a pattern compression and fusion module to decrease further the embedding size for the input of tpa, pg. 1480, left col., second para.) comprising a transformer-based architecture (We integrate the completeness matrix into the conventional self-attention mechanism (pg. 1479, left col., third para.); According to instant specification: “Once the data is preprocessed, it is passed to a latent transformer machine learning core 120. The machine learning core 120 employs advanced techniques such as self-attention mechanisms”(US20250363333 [0065])) configured to learn dependencies between elements of the fused vector (Pattern Compression and Fusion Layer in Fig. 3 produces a fused vector); and
generate a short-term forecast for the numerical time series data (transient spatial patterns capture short-term dynamics, pg. 1475, right col., first para.),
the short-term forecast comprising one or more predicted future values of the numerical time series data (persistent patterns common across different time series, enabling rapid pattern retrieval during inference, abstract), the predicted future values being generated by the machine learning core based on patterns identified by the machine learning core in historical data of the numerical time series data (a bidirectional-RNN model that incorporates information from both past ... timestamps, pg. 1476, left col., third para.) and based on the fused vector incorporating the textual data (Pattern Compression and Fusion Layer in Fig. 3 produces a fused vector).
Lai does not explicitly teach a compound large codeword model, comprising one or more computers with executable instructions that, when executed, cause the deep learning system to: allocate codewords to each data input, wherein codewords are mapped to a corresponding codebook, and wherein the codewords and their corresponding codebooks are adaptively updated to reflect incoming data inputs, a machine learning core comprising a transformer-based architecture
Harikumar teaches a deep learning system (In some implementations, a CNN is used to predict image tokens [0033]) for real-time time series forecasting (predict the next image tokens autoregressively (one token at a time in serial fashion [0041]) using a compound large codeword model (As used herein, in the context of this disclosure, a “codebook” is a visual dictionary of a set of visual words (or codewords) that represent one or more feature vectors of one or more images [0035]), comprising one or more computers with executable instructions that, when executed instructions that, when executed, cause the deep learning system to: (In some aspects, a computer-readable apparatus including a storage medium stores computer-readable and computer-executable instructions that are configured to, when executed by at least one processor apparatus [0102]),
allocate codewords to each data input, wherein codewords are mapped to a corresponding codebook (In some embodiments, a set of image tokens each having a unique integer value is obtained based on the codebook 502 a for the image encoder 510 and the codebook 502 b for the sketch encoder 520 [0070], Fig. 5. The Examiner notes a codebook contains codeword), and
wherein the codewords and their corresponding codebooks are adaptively updated to reflect incoming data inputs (codebook 502a is adaptively updated by incoming data inputs from 504a-n and codebook 502b is adaptively updated by incoming data inputs from 504b, Fig.5)
a machine learning core comprising a transformer-based architecture (Transformer models 106 and 604 are examples of the transformer model [0111])
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lai to incorporating the teachings of Harikumar for the benefits of using a transformer model which increases the resolution of the reconstructed data that results in the intricate patterns being created without loss of fidelity and at high efficiency (Harikumar [0099])
Regarding claim 2, Lai and Harikumar teaches the system of claim 1, wherein the transformer-based architecture of the machine learning core comprises a multi-head self-attention mechanism (The method for Completeness aware Spatial Attention follows a similar principle but applies self-attention across spatial data slices instead, pg. 1478, Fig. 4) and
a feed-forward network applied to each position of an input sequence (A Grouped Feedforward Network (GFFN) with two GFFN layers then takes the attentive embedding and performs light weight non-linear activation to obtain the final output: ... where G∗ is the number of groups and GFFN(· | G∗) is a GFFN layer, pg. 1479, right col., first para.).
Regarding claim 3, Lai and Harikumar teaches the system of claim 1, Harikumar teaches wherein the machine learning core further comprises uses a latent transformer (Transformer models 106 and 604 are examples of the transformer model [0111]; In various other embodiments, the encoder 404 and decoder 408 are transformer models [0051]) that operates directly on latent space vectors produced from the fused vector by a variational autoencoder encoder (In some implementations, the encoder is an autoencoder ... In some implementations, the autoencoder is a Vector Quantized Variational Autoencoder (VQVAE), which is configured to learn discrete (rather than continuous) latent representation of an image [0066]), the latent transformer being configured to process the latent space vectors (an unsupervised learning technique that uses a neural network to find non-linear latent representations for a given data distribution [0066]) without an embedding layer and without a positional encoding layer (The transformer model does not have an embedding layer or positional encoding layer).
Regarding claim 7, Lai teaches a method for real-time time series forecasting (One school of deep models is based on the common framework used for CTS (Correlated Time Series) forecasting, pg. 1476, left col., third para.) using a compound large codeword model (To achieve this, we design a novel Multi-view Learnable Codebook (MvLC) mechanism, pg. 1477, left col., fourth para.) comprising the steps of: receiving a plurality of data inputs (CTS Data Input, Fig. 1, pg. 1474) comprising at least textual data (Monday... Fig. 2, pg. 1475) and numerical time series data (12:00, 13:00 ..., Fig. 2: Multi-view patterns in CTS (pg. 1475); spatio-temporal and associate each with persistent and transient components, as shown in Figure 2. Temporal patterns encompass time-related features; persistent ones recur periodically (e.g., weekly traffic flow variations), pg. 1475, right col., first para.);
allocating codewords to each data input (PT-pattern Codebook (CBPT): Persistent Temporal Patterns (PT patterns) represent periodic information that can employ fixed pattern representation rules for all temporal segments, ... We assign aPT-pattern tensor vPT𝑡 to each timestamp 𝑡 and broadcast it across different time series, pg. 1477, right col., third para.),
wherein codewords are mapped to a corresponding codebook (MvLC Mechanism. The Learnable Codebook serves as the foundation for MvLC. It functions as a dictionary that maps input information to fixed-length pattern tensors, pg. 1477, right col., first full para.), and
wherein the codewords and their corresponding codebooks (PT-pattern Codebook, PS-pattern Codebook, PST-pattern Codebook, pg. 1477, right col.) are adaptively updated to reflect incoming data inputs by (To achieve this, we design a novel Multi-view Learnable Codebook (MvLC) mechanism, ... This mechanism learns and stores representations of three subtle views of persistent patterns during training, allowing for easy retrieval during imputation via periodic information updates, pg. 1477, left col. fourth para.);
analyzing the incoming data inputs to identify emerging patterns not represented by codewords of the corresponding codebook (An optimal window size of T = 24 hours aligns with the dataset’s daily periodicity, thereby optimizing the model’s capabilities, pg. 1482, left col., first para. The Examiner notes daily periodicity indicates daily recurring incoming data at regular, predictable intervals);
monitoring usage frequencies of the codewords of the corresponding codebook (After training, the codebooks of ppe are stored and used for subsequent imputation inference via table look-ups, pg. 1479, right col., third to the last para. The Examiner notes that using codebooks for subsequent imputation inference indicates the usage of frequency of the codebooks are monitored);
generating new codewords for the identified emerging patterns (Similar to the PT-pattern codebook, the Persistent Spatial Pattern (PS-pattern) codebook aims to capture patterns that remain constant over time but vary spatially. This could be geographical location or identity information of timeseries... These values are then broadcast across timestamps to achieve the PS-pattern, pg. 1477, right col., second to the last para. The Examiner notes new codewords are generated when codebook captures patterns that remain constant over time but vary spatially); and
pruning, from the corresponding codebook (Initially, we evaluate each codebook individually by systematically removing them, pg. 1480, right col., second to the last para.),
codewords whose monitored usage frequencies fall below a threshold (spatio-temporal attention requires 4096 times more computational and memory resources than does spatial attention alone, which is prohibitive. To achieve a resource-efficient model, we therefore limit the consideration of ST-patterns to the ppe phase, pg. 1479, left col., 4.3.2 Completeness-aware Attention);
fusing codewords of the textual data and codewords of the numerical time series data (PT-pattern Codebook, PS-pattern Codebook, PST-pattern Codebook are fused together (Fig. 3, pg. 1478); Monday...(day of the week); 12:00, 13:00...(time of day), geographical location (spatial pattern), Fig. 2: Multi-view patterns in CTS, pg. 1475), which are dissimilar data types (PT-pattern Codebook (CBPT), Persistent Temporal-pattern (PT-pattern)... For instance, information such as day-of-the-week (DoW) or time-of-the-day (ToD) is inherently periodic, pg. 1477, right col., third to last para.; PS-pattern Codebook ... aims to capture patterns that remain constant over time but vary spatially. This could be geographical location, pg. 1477, right col., third to last para.), together into a single codeword representation (To emphasize more significant persistent patterns, we concatenate all three views of persistent patterns along the feature embedding dimension, pg. 1478, left column, second para.) by extracting features from the codewords of each data type (This phase utilizes primarily the MvLC to extract persistent pat terns from three distinct views: the Persistent Temporal Pattern (PT-pattern), the Persistent Spatial Pattern (PS-pattern), and the Persistent Spatio-temporal Pattern (PST-pattern) through learnable codebooks, pg. 1477, left col., last para.), capturing interactions between the extracted features (pattern learning architecture that represents are thinking of the pattern extraction process. CTS typically encompasses two types of patterns: persistent patterns, capturing how data points evolve in relation to static information, such as temporal cycles and the geographic location of time series, and transient patterns, describing fluctuations at specific timestamps and semantic spatial correlations, pg. 1477. Left col., second full para.), and
combining the extracted features and interactions through a fusion layer to produce a fused vector (Pattern Compression and Fusion Layer in Fig. 3 produces a fused vector);
processing the fused vector representing the single codeword representation through a machine learning core (ReCTSi integrates a pattern compression and fusion module to decrease further the embedding size for the input of tpa, pg. 1480, left col., second para.) comprising a latent transformer architecture (We integrate the completeness matrix into the conventional self-attention mechanism (pg. 1479, left col., third para.); According to instant specification: “Once the data is preprocessed, it is passed to a latent transformer machine learning core 120. The machine learning core 120 employs advanced techniques such as self-attention mechanisms”(US20250363333 [0065]))
and generating a short-term forecast for the numerical time series data (transient spatial patterns capture short-term dynamics, pg. 1475, right col., first para.),
the short-term forecast comprising one or more predicted future values of the numerical time series data (persistent patterns common across different time series, enabling rapid pattern retrieval during inference, abstract),
the predicted future values being generated by the machine learning core based on patterns identified by the machine learning core in historical data of the numerical time series data (a bidirectional-RNN model that incorporates information from both past ... timestamps, pg. 1476, left col., third para.) and based on the fused vector incorporating the textual data (Pattern Compression and Fusion Layer in Fig. 3 produces a fused vector).
Lai does not explicitly teach a compound large codeword model, allocating codewords to each data input, wherein codewords are mapped to a corresponding codebook, and wherein the codewords and their corresponding codebooks are adaptively updated to reflect incoming data inputs, a machine learning core comprising a latent transformer architecture that operates on latent space vectors produced from the fused vector by variational autoencoder encoder, without an embedding layer and without a positional encoding layer; and generating a short-term forecast for the numerical time series data, the short-term forecast comprising one or more predicted future values of the numerical time series data, the predicted future values being generated by the machine learning core based on patterns identified by the machine learning core in historical data of the numerical time series data and based on the fused vector incorporating the textual data.
Harikumar teaches a deep learning system (In some implementations, a CNN is used to predict image tokens [0033]) for real-time time series forecasting (predict the next image tokens autoregressively (one token at a time in serial fashion [0041]) using a compound large codeword model (As used herein, in the context of this disclosure, a “codebook” is a visual dictionary of a set of visual words (or codewords) that represent one or more feature vectors of one or more images [0035]),
allocate codewords to each data input, wherein codewords are mapped to a corresponding codebook (In some embodiments, a set of image tokens each having a unique integer value is obtained based on the codebook 502 a for the image encoder 510 and the codebook 502 b for the sketch encoder 520 [0070], Fig. 5. The Examiner notes a codebook contains codeword), and
wherein the codewords and their corresponding codebooks are adaptively updated to reflect incoming data inputs (codebook 502a is adaptively updated by incoming data inputs from 504a-n and codebook 502b is adaptively updated by incoming data inputs from 504b, Fig.5)
a machine learning core comprising a latent transformer architecture Transformer models 106 and 604 are examples of the transformer model [0111] ; In various other embodiments, the encoder 404 and decoder 408 are transformer models [0051]) that operates on latent space vectors produced from the fused vector by variational autoencoder encoder (In some implementations, the encoder is an autoencoder ... In some implementations, the autoencoder is a Vector Quantized Variational Autoencoder (VQVAE), which is configured to learn discrete (rather than continuous) latent representation of an image [0066]), without an embedding layer and without a positional encoding layer (The transformer model does not have an embedding layer or positional encoding layer);
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lai to incorporating the teachings of Harikumar for the benefits of using a transformer model which increases the resolution of the reconstructed data that results in the intricate patterns being created without loss of fidelity and at high efficiency (Harikumar [0099])
Regarding claim 8, Lai and Harikumar teaches the method of claim 7, wherein the latent transformer architecture of the machine learning core comprises a multi-head self-attention mechanism (The method for Completeness aware Spatial Attention follows a similar principle but applies self-attention across spatial data slices instead, pg. 1478, Fig. 4) and
a feed-forward network applied to each position of an input sequence of the latent space(A Grouped Feedforward Network (GFFN) with two GFFN layers then takes the attentive embedding and performs light weight non-linear activation to obtain the final output: ... where G∗ is the number of groups and GFFN(· | G∗) is a GFFN layer, pg. 1479, right col., first para.).
5. Claims 5 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. ("ReCTSi: resource-efficient correlated time series imputation via decoupled pattern learning and completeness-aware attentions." Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2024) in view of Harikumar et al. (US20230419551) and further in view of Guo et al. ("MSMC-TTS: Multi-stage multi-codebook VQ-VAE based neural TTS." IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2023): 1811-1824)
Regarding claim 5, Lai and Harikumar teaches the system of claim 1, Lai teaches wherein the short-term forecast (transient spatial patterns capture short-term dynamics, pg. 1475, right col., first para.) is generated by creating, for the numerical time series data (persistent patterns common across different time series, enabling rapid pattern retrieval during inference, abstract),
an input vector comprising a truncated portion of historical data points followed by a sequence of placeholder values (As exemplified in Figure 4, the percentages of missing values in the first and fourth timestamps (columns) are 0.33, much
lower than in the other two timestamps (pg. 1478, right col., last para.); Given the mask matrix M associated with an input CTS X, the temporal completeness rate of a timestamp 𝑡, defined as its percentage of missing values, is computed as follows:
PNG
media_image1.png
48
414
media_image1.png
Greyscale
where M𝑛,𝑡 is the element of M at position (𝑛,𝑡). The Examiner notes that temporal completeness rate of a timestamp 𝑡, defined as its percentage of missing values
associated with the input vector is the truncated portion of historical data points),
Lai and Harikumar does not explicitly teach processing the input vector through a variational autoencoder encoder to obtain a latent space representation, and processing the latent space representation through the machine learning core to populate the placeholder values with the predicted future values.
Guo teaches processing the input vector through a variational autoencoder encoder to obtain a latent space representation, and processing the latent space representation through the machine learning core (In each encoder, the input sequence is processed by a projection layer, and added with position encodings, then fed to 4 FeedForward Transformer blocks, pg. 1816, right col., second para.) to populate the placeholder values with the predicted future values (the model can be given the predicted sequence ˆz(i) or the ground-truth sequence z(i) at i-th stage, pg. 1820, right col., second to the last para.).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lai and Harikumar to incorporating the teachings of Guo for the benefits of lowering modeling complexity and data size requirements, preserving excellent performance even with fewer model parameters or training data (Guo, abstract)
Regarding claim 11, claim 11 is similar to claim 5. It is rejected in the same manner and reasoning applying.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MORIAM MOSUNMOLA GODO whose telephone number is (571)272-8670. The examiner can normally be reached Monday-Friday 8am-5pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle T Bechtold can be reached on (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/M.G./Examiner, Art Unit 2148
/MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148