DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1, 11, and 20 recite the limitation "wherein the patch-level predictions…" There is insufficient antecedent basis for this limitation in the claim, as there is no prior mention of patch-level predictions. Thus, claims 1, 11, and 20 (and associated dependent claims) are rejected under 35 U.S.C. 112.
Claims 1, 11, and 20 recite that “the patch-level predictions are generated from outputs of the gated MLP mixing”, indicating that the ‘patch-level predictions’ are not necessarily directly generated by the gated MLP mixing, but are generated based on outputs of the gated MLP mixing. The claims further recite “patch-level predictions generated by the gated MLP mixing,” which is inconsistent with the previous mention of how the ‘patch-level predictions’ are generated, which renders the scope of the claim indefinite. Thus, claims 1, 11, and 20 (and associated dependent claims) are rejected under 35 U.S.C. 112. For examination purposes, the ‘patch-level predictions’ are being interpreted as predictions generated from outputs of the gated MLP mixing, not generated directly by the gated MLP mixing.
The following is a quotation of 35 U.S.C. 112(d):
(d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph:
Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
Claims 2, 12 rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. Independent claims 1 and 11 already state that the MLP mixing is channel independent. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Step 1
According to the first part of the analysis, in the instant case, claims 1-4, 6-10 are directed to a method, claims 11-2, 14-20 are directed to an apparatus. Each of these claims fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter).
Regarding claim 1
Step 2A Prong One
applying gated multilayer perceptron (MLP) mixing across different directions of the patched input time-series, … capturing local and global and interrelated correlations across the plurality of patches and within the plurality of patches
(These steps for applying a mathematical model [gated MLP mixing] to capture correlations in data is a mathematical concept)
and applying a patch-time aggregated hierarchy to guide lowest-level predictions based on aggregated hierarchy signals at a patch-level,
(This step for guiding predictions based on signals is a mental process)
wherein the patch-level predictions are generated from outputs of the gated MLP mixing after capturing the local and global and interrelated correlations across and within the plurality of patches,
(This step for generating predictions from / based on outputs of the gated MLP mixing is a mental process)
Step 2A Prong Two
A computer-implemented method comprising:
(This step for performing the methods on a generic computer is considered mere instructions to apply an exception. See MPEP § 2106.05(f))
segmenting a time-series dataset from a plurality of sensors into a plurality of patches that are overlapping or non-overlapping;
(This step for pre-processing data is insignificant extra-solution activity. See MPEP § 2106.05(g))
wherein the MLP mixing is channel independent and shares weights across channels and mixes across the plurality of patches and across features of the plurality of patches; capturing … correlations … by alternately mixing along patch and feature dimensions using the gated MLP mixing;
These additional elements integrate the mathematical concept of “applying gated multilayer perceptron (MLP) mixing across different directions of the patched input time-series and the mental process of capturing local and global and interrelated correlations across the plurality of patches and within the plurality of patches” as it improves the functionality of the technology ([0030] “Mixing in a channel independent way across different directions means alternating with respect to patches and features. By mixing in this manner, aspects of this disclosure are able to capture local and global interrelated correlations across and within patches and features.”)
and wherein the patch-time aggregated hierarchy comprises generating hierarchical aggregated forecasts from the patch-level predictions generated by the gated MLP mixing and reconciling the generated hierarchical aggregated forecasts.
(These steps for generating and reconciling forecasts based on the aforementioned predictions is insignificant extra-solution activity. See MPEP § 2106.05(g))
Step 2B
The claim recites mental processes while the additional elements are considered mere instructions to apply an exception and insignificant extra-solution activity, however, the additional element of “generating hierarchical aggregated forecasts from the patch-level predictions generated by the gated MLP mixing and reconciling the generated hierarchical aggregated forecasts” is extra-solution activity which is not well understood, routine, and conventional activity, thus, the claim includes additional elements that are sufficient to amount to significantly more than the remaining judicial exceptions, and the claim is not rejected under 101.
Claim 11 is an apparatus claim substantially similar to claim 1, and is analyzed under 101 using the same reasoning.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 8-14, 17-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhengzhong Tu et al. (hereinafter Tu) (“MAXIM: Multi-Axis MLP for Image Processing”, 04/09/2022) in view of Maja Rudolph et al. (hereinafter Maja) (JP 2023010698 A, 01/20/2023) further in view of Yunhao Zhang et al. (hereinafter Zhang) (“CROSSFORMER: TRANSFORMER UTILIZING CROSS DIMENSION DEPENDENCY FOR MULTIVARIATE TIME SERIES FORECASTING”, 03/14/2023) further in view of Davide Burba et al. (hereinafter Burba) (“A Trainable Reconciliation Method for Hierarchical Time-Series”, 01/05/2021), further in view of Huang Jinmiao et al. (hereinafter Huang) (US 20230335118 A1, 10/19/2023) further in view of Ilya Tolstikhin et al. (hereinafter Tolstikhin) (“MLP-Mixer: An all-MLP Architecture for Vision”, 2021).
Regarding claim 1, Tu teaches;
applying gated multilayer perceptron (MLP) mixing across different directions of the patched input time-series;
PNG
media_image1.png
509
1451
media_image1.png
Greyscale
([pg. 4] Figure 3. Multi-axis gated MLP block (best viewed in color). The input is first projected to a [6; 4;C] feature, then split into two heads. In the local branch, the half head is blocked into 32 non-overlapping [2; 2;C=2] patches, while we grid the other half using a 2x2 grid in the global branch. We only apply the gMLP block [50] (illustrated in the right gMLP Block) on a single axis of each branch - the 2nd axis for the local branch and the 1st axis for the global branch, while shared along the other spatial dimensions. The gMLP operators, which run in parallel, correspond to local and global (dilated) attended regions, as illustrated with different colors (i.e., the same color are spatially mixed using the gMLP operator). Our proposed block expresses both global and local receptive fields on arbitrary input resolutions.)
NOTE: Teaches applying gated MLP mixing (mixed using gMLP, which is gated MLP) across different directions of the input (1st and 2nd axis, as shown in fig. 3 above). It would be obvious for this patched input to be patched time-series data, further explained below.
capturing local and global and interrelated correlations across the plurality of patches and within the plurality of patches;
PNG
media_image2.png
520
1454
media_image2.png
Greyscale
([pg.4, fig. 3] Multi-axis gated MLP block (best viewed in color). The input is first projected to a [6; 4;C] feature, then split into two heads. In the local branch, the half head is blocked into 3 2 non-overlapping [2; 2;C=2] patches, while we grid the other half using a 2 2 grid in the global branch. We only apply the gMLP block [50] (illustrated in the right gMLP Block) on a single axis of each branch - the 2nd axis for the local branch and the 1st axis for the global branch, while shared along the other spatial dimensions. The gMLP operators, which run in parallel, correspond to local and global (dilated) attended regions, as illustrated with different colors (i.e., the same color are spatially mixed using the gMLP operator). Our proposed block expresses both global and local receptive fields on arbitrary input resolutions.)
NOTE: The local branch restricts mixing to local windows while the global branch mixes across global spatial areas (across and within the plurality of patches), the two outputs are then concatenated {see fig. 3}, capturing local and global interrelations respectively.
wherein the
[pg. 4]
PNG
media_image3.png
115
552
media_image3.png
Greyscale
NOTE: Tu teaches MAXIM capturing the local and global and interrelated correlations across and within the plurality of patches using gated MLP mixing (as previously taught), and teaches the output of MAXIM (i.e., after capturing the global and local correlations) being predictions at different scales / levels.
Tu fails to teach but Maja teaches;
segmenting a time-series dataset from a plurality of sensors into a plurality of patches that are overlapping or non-overlapping
([Abstract] receiving time series data being grouped in patches)
NOTE: Teaches segmenting a time-series dataset into a plurality of patches.
([pg. 2] The technology disclosed in the present application can be applied to time-series imaging using other sensors such as radio electromagnetic wave antennas and sound collecting microphones, for example.)
NOTE: The aforementioned time-series dataset can be derived from a plurality of sensors.
OBVIOUSNESS TO COMBINE MAJA WITH TU:
Maja and Tu are both analogous art to the present disclosure as Maja discloses a system which patches time-series data from sensors while Tu discloses a process using gated MLP mixing.
Tu teaches a gated MLP-mixer architecture operating on data patches, while Maja teaches patching time-series sensor data for further processing.
Maja additionally states;
([pg. 2] Images may be multi-dimensional in that they may include components of time, space, intensity, intensity, or other properties. For example, the images may include time series images.)
NOTE: Maja discloses that adding a time dimension to image data is a known technique to allow the data to represent patterns over time. The disclosure of Tu pertains to image data.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to collect and segment data from a plurality of sensors with a time series component into patches (as taught by Maja) to use as input for a system utilizing gated MLP mixing (taught by Tu) to capture temporal trends in data.
Tu and Maja fail to teach but Zhang teaches;
applying a patch-time aggregated hierarchy … patch-level predictions … and wherein the patch-time aggregated hierarchy comprises generating hierarchical aggregated forecasts from the patch-level predictions
[pg. 3]
PNG
media_image4.png
603
620
media_image4.png
Greyscale
NOTE: Zhang teaches a hierarchical encoder-decoder architecture which generates predictions / forecasts for each time-series segment (patch) scale and adds them to generate hierarchical aggregated forecasts. Thus, Zhang teaches a patch-time aggregated hierarchy (a patch-time hierarchy which utilizes aggregation) comprising generating hierarchical aggregated forecasts from the patch-level predictions (adds [aggregates] forecasts [predictions] from different segment [patch]-levels of the hierarchy to generate a final hierarchical aggregated prediction [forecast]).
OBVIOUSNESS TO COMBINE ZHANG WITH TU, MAJA:
Zhang is analogous art to the present disclosure as it pertains to time-series forecasting using temporal patches and an aggregated hierarchy architecture.
Tu teaches generating predictions based on patched data at multiple scales using a gated MLP-Mixer architecture, Maja teaches patching time-series sensor data, and Zhang teaches generating time-series patch-level predictions / forecasts to be utilized in an aggregated hierarchy forecasting architecture.
Zhang states that the hierarchical architecture is used to make a final prediction using predictions at multiple temporal scales, rather than only one fixed patch / scale, predictably improving time-series forecasting accuracy by incorporating multi-scale temporal information;
([pg. 5] Architecture of the Hierarchical Encoder-Decoder in Crossformer with 3 encoder layers… Exploring different scales, the decoder (right) makes the final prediction by forecasting at each scale and adding them up.)
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to use the hierarchical forecasting architecture of Zhang within the system taught by the combination of Tu and Maja, to make a final prediction using predictions at multiple patch levels, rather than only one fixed patch level, to improve overall prediction accuracy by incorporating multi-scale temporal information.
From this, the combination of Tu and Zhang reasonably teach;
wherein the patch-level predictions [Zhang] are generated from outputs of the gated MLP mixing after capturing the local and global and interrelated correlations across and within the plurality of patches [Tu]
Tu, Maja, and Zhang fail to teach but Burba teaches;
and applying a
PNG
media_image5.png
112
629
media_image5.png
Greyscale
NOTE: Fig 1 (above) Shows the time series hierarchy, where each level is an aggregate of the time-series segments from the layer below it. Thus, Burba teaches a time-series aggregated hierarchy.
PNG
media_image6.png
364
939
media_image6.png
Greyscale
([pg.4, section 3] The encoder maps the input predictions to the bottom level reconciled predictions, and the decoder takes the latter as input and reconstructs the predictions at all levels. A representation of our method is given in Figure 2.)
NOTE: The encoder maps all predictions to bottom level reconciled predictions, which is used to reconstruct predictions at all levels. Thus, Burba teaches guiding lowest level predictions (reconstructs predictions at all levels, including lowest level predictions) based on aggregated hierarchy signals at a level (each level contributes to the reconstructed reconciled predictions)
and wherein the
([pg. 3] producing forecasts only for the bottom level of the hierarchy and then aggregating them to the upper levels.)
NOTE: Burba discloses a hierarchical forecasting approach which produces forecasts for lower levels of the hierarchy then aggregates those forecasts to upper levels. Accordingly, Burba teaches generating hierarchical aggregated forecasts from lower-level predictions / forecasts, where upper-level forecasts are generated by aggregating lower-level forecasts according to the hierarchy.
and reconciling the generated hierarchical aggregated forecasts.
PNG
media_image6.png
364
939
media_image6.png
Greyscale
NOTE: Burba teaches reconciling predictions at all hierarchy levels, including any generated hierarchical aggregated forecasts / predictions.
OBVIOUSNESS TO COMBINE BURBA WITH TU, MAJA, AND ZHANG:
Burba is analogous art to the present disclosure as it relates to time series forecasting using an aggregated hierarchy structure with reconciliation.
Tu teaches a gated MLP mixer architecture for using patched data to generate predictions at different scales, Maja teaches patching time-series sensor data, Zhang teaches a patch-time aggregated hierarchy architecture, and Burba teaches a time-series aggregated hierarchy architecture with reconciliation.
Burba further states that hierarchical levels typically do not add up properly because of hierarchical restraints, so a reconciliation is needed to ensure the hierarchy is satisfied, e.g., parent forecasts match the sum of child forecasts;
([Abstract] independent forecasts typically do not add up properly because of the hierarchical constraints, so a reconciliation step is needed.)
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the patch-time aggregated hierarchy architecture of Zhang to include the reconciliation step of Burba to ensure the hierarchy is satisfied, predictably improving coherence and consistency across the hierarchy.
From this, the combination of Tu, Zhang, and Burba reasonably teaches;
and wherein the patch-time aggregated hierarchy comprises generating hierarchical aggregated forecasts [Zhang and Burba] from the patch-level predictions [Zhang] generated by the gated MLP mixing [Tu] and reconciling the generated hierarchical aggregated forecasts. [Burba]
Tu, Maja, Zhang and Burba fail to teach but Huang teaches;
wherein the MLP mixing … mixes across the plurality of patches and across features of the plurality of patches
([0094] The basic idea behind MLP-Mixer is to use multiple layers of mixer blocks, each consisting of two separate operations: channel mixing and spatial mixing. In the channel mixing operation, a multi-layer perceptron (MLP) is applied to the channels of each patch independently, allowing for non-linear transformations of the patch features. In the spatial mixing operation, a global average pooling is performed over the patch dimensions)
NOTE: Huang teaches an MLP mixer which includes both spatial and channel mixing, where spatial mixing is performed over the patch dimension, and channel mixing mixes within each patch across its feature dimension.
capturing local and global and interrelated correlations across the plurality of patches and within the plurality of patches by alternately mixing along patch and feature dimensions using the
([0094] In the channel mixing operation, a multi-layer perceptron (MLP) is applied to the channels of each patch independently, allowing for non-linear transformations of the patch features. In the spatial mixing operation, a global average pooling is performed over the patch dimensions … [0095] By alternating these two operations, MLP-Mixer can capture both local and global features])
NOTE: Huang teaches capturing local and global and interrelated correlations across the plurality of patches and within the plurality of patches (capture both local and global features of the patches) by alternately mixing (the MLP-Mixer alternates between spatial and channel mixing) along patch (spatial mixing is performed over the patch dimension) and feature dimensions (channel mixing mixes within each patch - across its feature dimension) using the MLP mixing.
OBVIOUSNESS TO COMBINE HUANG WITH TU, MAJA, ZHANG AND BURBA:
Huang is analogous art to the present disclosure as it pertains to MLP mixing.
Tu teaches using a gated MLP mixing architecture on patched data to capture local and global correlations, Maja teaches patching time-series sensor data, and Huang teaches alternately mixing along patch and feature dimensions to capture local and global correlations.
Huang further states that by alternating these mixing operations, the MLP-Mixer can capture global features and local features while also allowing for non-linear transformations of the feature representations.
([0095] By alternating these two operations, MLP-Mixer can capture both local and global features of the image, while also allowing for non-linear transformations of the feature representations)
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the gated-MLP mixing of Tu to alternatively mix along patch and feature dimensions (as taught by Huang) to additionally allow for non-linear transformations of the feature representations while capturing global and local correlations.
From this, the combination of Tu and Huang reasonably teaches;
capturing local and global and interrelated correlations across the plurality of patches and within the plurality of patches [Tu and Huang] by alternately mixing along patch and feature dimensions [Huang] using the gated MLP mixing [Tu]
Tu, Maja, Zhang, Burba, and Huang fail to teach but Tolstikhin teaches;
wherein the MLP mixing is channel independent and shares weights across channels
([pg. 2] The token-mixing MLPs … operate on each channel independently … [pg. 3] token-mixing MLPs in Mixer that share the same kernel (of full receptive field) for all of the channels.)
OBVIOUSNESS TO COMBINE TOLSTIKHIN WITH TU, MAJA, ZHANG, BURBA, HUANG:
Tolstikhin is analogous art to the present disclosure as it pertains to an MLP-Mixer architecture.
Tu provides a gated MLP-mixer operating on patches, Maja teaches patching time-series data, Huang teaches alternating mixing on the patch and feature dimension, and Tolstikhin teaches mixing which is channel independent and shares weights across channels.
Tolstikhin additionally states that by sharing weights across channels, the disclosed mixer method leads to significant memory savings;
([pg. 3] token-mixing MLPs in Mixer that share the same kernel (of full receptive field) for all of the channels. The parameter tying prevents the architecture from growing too fast when increasing the hidden dimension C or the sequence length S and leads to significant memory savings.)
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the MLP mixer architecture taught by the combination of Tu and Zhang to include the channel independent weight sharing technique disclosed by Tolstikhin, to allow for significant memory savings.
Regarding claim 2, Tu, Maja, Zhang, and Burba fail to teach but Huang teaches;
The computer-implemented method of claim 1, wherein the MLP mixing is channel independent.
(channel independent mixing means alternating with respect to patches and features, as defined in the spec [0032])
([0094] The basic idea behind MLP-Mixer is to use multiple layers of mixer blocks, each consisting of two separate operations: channel mixing and spatial mixing. In the channel mixing operation, a multi-layer perceptron (MLP) is applied to the channels of each patch independently, allowing for non-linear transformations of the patch features. In the spatial mixing operation, a global average pooling is performed over the patch dimensions, followed by another MLP applied to the resulting global feature map.)
OBVIOUSNESS:
Using the same reasoning from claim 1, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the gated-MLP mixing of the system taught by the combination of Tu, Maja, and Burba, to alternatively mix along patch and feature dimensions (as taught by Huang) to additionally allow for non-linear transformations of the feature representations while capturing global and local correlations.
Regarding claim 3, Tu, Maja, Zhang, and Burba fail to teach but Huang teaches;
wherein the MLP mixing uses layers that are stacked in linear fashion.
PNG
media_image7.png
719
757
media_image7.png
Greyscale
NOTE: In fig. 6A {above} the MLP mixing uses layers that are stacked in a linear fashion (feature mixing layer -> time mixing layer -> etc.)
Additionally, stacking mixing layers in a linear fashion would allow the system to alternate between different mixing modes, allowing the mixing to generate insights across different dimensions of the data.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to make the MLP mixing of the system of claim 1 (taught by Tu in view of Maja and Burba) stack the MLP mixing layers in a linear fashion to allow the system to alternate between different mixing modes and generate insights across different dimensions of the data.
Regarding claim 4, Tu teaches;
MLP mixing uses layers
([pg. 3] we insert the multi-axis gated MLP block (MAB) into each encoder, decoder, and bottleneck (Fig.2b), with a residual channel attention block … stacked subsequently.)
NOTE: MAXIM teaches MLP mixing using layers because its encoder / decoder / bottleneck layers include repeated multi-axis gated MLP blocks, which mix local and global information.
Tu, Maja, and Zhang fail to teach but Burba teaches;
PNG
media_image5.png
112
629
media_image5.png
Greyscale
OBVIOUSNESS:
([pg.1, section 1] A hierarchical time-series is a collection of time-varying observations organized in a hierarchical structure. The problem of forecasting hierarchical time-series often appears in business and economics, where time-varying quantities need to be predicted at different granularity levels.)
NOTE: Representing time-series data using layers that are chained in a patch length context aware hierarchy fashion allows for time-varying quantities to be predicted at different granularity levels, which is a common need for this type of data [see introduction of Burba].
Hierarchical time-series forecasting commonly requires capturing dependencies at different temporal scales, and applying mixing operations across representations generated at each hierarchical level allows information from multiple patches to interact, thereby improving context awareness and predictive performance. Such a combination represents a predictable use of known network design techniques that combines hierarchical feature representation with mixing layers.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, for the MLP mixing of Tu to be applied within the hierarchical time-series segment structure of Burba to improve modeling of temporal relationships across multiple granularities.
Claim 5 has been cancelled.
Regarding claim 8, Tu fails to teach but Maja teaches;
Values of the plurality of sensors
[Using the same teaching from claim 1]
Tu and Maja fail to teach but Zhang teaches;
further comprising a downstream task of forecasting values
([pg. 7] Predictions of all the layers are summed to obtain the final forecasting)
Using the same reasoning from claim 1, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to use the hierarchical forecasting architecture of Zhang within the system taught by the combination of Tu and Maja, to make a final prediction using predictions at multiple patch levels, rather than only one fixed patch level, to improve overall prediction accuracy by incorporating multi-scale temporal information.
Regarding claim 9, Tu fails to teach but Maja teaches;
Values of the plurality of sensors
[Using the same teaching from claim 1]
Tu, Maja, and Zhang fail to teach but Burba teaches;
further comprising a downstream task of executing regression analysis regarding values
([pg.5, section 4.1] To train the forecasting models we consider two alternative strategies, … For the individual strategy, we consider linear autoregressive models (AR(p)), taking the lagged time-series values as input.)
NOTE: Burba teaches a task of training the forecasting models using regression analysis (linear autoregressive models) on time-series values. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to use a downstream task of executing regression analysis to train the forecasting models.
Regarding claim 10, Tu fails to teach but Maja teaches;
values of the plurality of sensors
(Using the same reasoning from claim 1)
Tu, Maja, Zhang, Burba, and Huang fail to teach but Tolstikhin teaches;
a downstream task of classifying values
([pg. 3] We evaluate the performance of MLP-Mixer models, pre-trained with medium- to large-scale datasets, on a range of small and mid-sized downstream classification tasks.)
Tolstikhin further states that their MLP-Mixer architecture attains competitive classification performance while maintaining acceptable resource usage;
([Abstract] MLP-Mixer attains competitive scores on image classification benchmarks, with pre-training and inference cost comparable to state-of-the-art models.)
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize the MLP-Mixer system using at least the methods disclosed by Tolstikhin to perform a downstream classification task to achieve competitive classification results and with acceptable resource usage.
Claim 11 is an apparatus claim directly corresponding to method claim 1 except with an additional limitation, which is taught by Maja;
A system comprising: a processor; and a memory in communication with the processor, the memory containing instructions that, when executed by the processor, cause the processor to:
PNG
media_image8.png
516
409
media_image8.png
Greyscale
([pg. 6] FIG. 5 is a block diagram of an electronic computer system suitable for implementing the systems disclosed herein or for performing the methods disclosed herein.)
NOTE: Teaches a system comprising: a processor; and a memory in communication with the processor, the memory containing instructions that, when executed by the processor, cause the processor to perform the methods of the disclosure.
Claim 12 is an apparatus claim directly corresponding to claim 2, and is rejected using the same reasoning.
Claim 13 has been cancelled.
Claim 14 is an apparatus claim directly corresponding to claim 4, and is rejected using the same reasoning.
Claim 17 is an apparatus claim directly corresponding to claim 8, and is rejected using the same reasoning.
Claim 18 is an apparatus claim directly corresponding to claim 9, and is rejected using the same reasoning.
Claim 19 is an apparatus claim directly corresponding to claim 10, and is rejected using the same reasoning.
Claim 20 is an apparatus claim directly corresponding to method claim 1 except with an additional limitation, which is taught by Maja;
A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to:
([pg. 15] Program code embodying the algorithms and/or methodologies described herein may be distributed separately or collectively in a variety of different forms of program products. This program code may be distributed using a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to perform aspects of one or more embodiments. Computer-readable storage media that are non-transitory in nature include both volatile and nonvolatile implemented by any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data.)
NOTE: Teaches a computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform the methods of the disclosure.
Claim(s) 6, 7, 15-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tu (“MAXIM: Multi-Axis MLP for Image Processing”, 04/09/2022) in view of Maja (JP 2023010698 A, 01/20/2023) further in view of Zhang (“CROSSFORMER: TRANSFORMER UTILIZING CROSS DIMENSION DEPENDENCY FOR MULTIVARIATE TIME SERIES FORECASTING”, 03/14/2023) further in view of Burba (“A Trainable Reconciliation Method for Hierarchical Time-Series”, 01/05/2021), further in view of Huang (US 20230335118 A1, 10/19/2023) further in view of Tolstikhin (“MLP-Mixer: An all-MLP Architecture for Vision,” 2021) as applied to claims 1 and 11 above, further in view of Krishna Kumar Singh et al. (hereinafter Krishna) (“Hide-and-Seek: Forcing a Network to be Meticulous for Weakly-supervised Object and Action Localization”, 12/23/2017).
Regarding claim 6, Tu, Maja, Zhang, Burba, Huang, and Tolstikhin fail to teach but Krishna teaches;
further comprising a pretraining task of masking random patches.
([Abstract] Our key idea is to hide patches in a training image randomly, forcing the network to seek other relevant parts when the most discriminative part is hidden.)
NOTE: Teaches randomly masking patches.
OBVIOUS TO COMBINE KRISHNA:
Krishna is analogous art to the present disclosures as it pertains to machine learning architectures.
Tu teaches using patched input data to generate predictions, while Krishna pertains to masking random patches of inputs.
Krishna further states:
([pg.1, fig.1 caption] Main idea. (Top row) A network tends to focus on the most discriminative parts of an image (e.g., face of the dog) for classification. (Bottom row) By hiding images patches randomly, we can force the network to focus on other relevant object parts in order to correctly classify the image as ’dog’.)
NOTE: This excerpt discloses that random masking allows the network to accurately predict even when data is missing.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate random masking (as taught by Krishna) into the architecture of Tu to allow the system to perform well even when data is missing.
Regarding claim 7, Tu, Maja, Zhang, Burba, Huang, and Tolstikhin fail to teach but Krishna teaches;
further comprising reconstructing the masked random patches.
([pg.4, paragraph 2] We hide patches only during training. During testing, the full image—without any patches hidden—is given as input to the network;)
NOTE: Teaches reconstructing (by revealing the masked patches in the original image) the masked random patches.
Krishna additionally states;
([pg. 4, paragraph 2] Since the network has learned to focus on multiple relevant parts during training, it is not necessary to hide any patches during testing.)
NOTE: There is no reason for the patches to be masked during testing, so it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to reconstruct the masked patches during testing to accurately test the model.
Claim 15 is an apparatus claim directly corresponding to claim 6, and is rejected for the same reasons.
Claim 16 is an apparatus claim directly corresponding to claim 7, and is rejected for the same reasons.
Response to Arguments
In review of Applicant’s amendments, filed 05/29/2026, the rejections from the previous office action made under 35 U.S.C 112 have been withdrawn. However, amendments to the claims have necessitated new rejections under 35 U.S.C 112, as outlined above.
Applicant's arguments filed 05/29/2026 regarding the 35 U.S.C. 101 rejections have been fully considered and are mostly un-persuasive.
Starting on page 3, regarding step 2A prong one of the 101 analysis, the applicant remarks;
Amended claim 1 is directed to a specific improvement in computer technology, not an abstract idea. Amended claim focuses on patch-based segmentation (overlapping or non-overlapping), channel-independent weight sharing architecture, directional MLP mixing across patch and feature dimensions, and hierarchical reconciliation for producing hierarchy-consistent forecasts. These operations cannot be performed mentally or with pen and paper.
Applicant suggests that the listed elements of claim 1 are not an abstract idea because the operations cannot be performed mentally or with pen and paper. Examiner agrees that limitations corresponding to the patch-based segmentation and hierarchical reconciliation are not abstract ideas, as reflected in the new 101 analysis. However, using a mathematical model (channel-independent MLP mixing) to capture correlations across data is a mathematical concept / abstract idea.
Amended claim 1 also recites additional limitations that are considered abstract ideas which are not addressed in the arguments, including;
applying a patch-time aggregated hierarchy to guide lowest-level predictions based on aggregated hierarchy signals at a patch-level
Guiding generic predictions (e.g., predictions generated mentally by a human) based on signals is a mental process. For example, the way the claim is structured, a human could look at signals generated by the claimed aggregated hierarchy, and use that information to guide mental predictions.
wherein the patch-level predictions are generated from outputs of the gated MLP mixing after capturing the local and global and interrelated correlations across and within the plurality of patches
Generating predictions from (i.e. based on) outputs is a mental process, using similar reasoning.
The applicant further recites the elements of amended claim 1 and portions of their spec that allegedly indicate that amended claim 1 reflects the technical solution to a technical problem, which is not relevant for step 2A prong one of the analyses for 35 U.S.C. 101.
Regarding step 2A prong two, starting on page 5, the applicant states that the amended claim no longer recites generic forecasting or correlation analysis with quote from the spec. Even assuming arguendo that this is true, this does not indicate that the claim as a whole integrates the abstract ideas into a practical application.
The applicant further includes a quote from the spec that allegedly identifies a technical benefit; “The architecture also includes attaching and tuning online hierarchical reconciliation head to the MLP-Mixer backbone. In this way, the architecture converts the learning capability of simple MLP structures to outperform complex transformer models with significantly less computing resources.” However, the claim language does not include ‘attaching and tuning an online hierarchical reconciliation head to an MLP-Mixer backbone.’ Thus, language from the spec which is not in the claim language cannot integrate the abstract ideas into a practical application.
However, the amended claims 1 and 11 include a new limitation which allows the claim as a whole to amount to significantly more, as reflected by the 101 analyses in this office action.
In consideration of this conclusion, the claims are deemed subject-matter eligible, and, thus, the rejections under 35 U.S.C. 101 have been removed.
Applicant's arguments filed 05/29/2026 regarding the 35 U.S.C. 103 rejections have been fully considered but they are not persuasive.
Starting on page 6, the applicant remarks;
The cited references do not teach or suggest the feature of amended independent claim 1 reciting, "applying gated multilayer perceptron (MLP) mixing across different directions of the patched input time-series, wherein the MLP mixing is channel independent and shares weights across channels and mixes across the plurality of patches and across features of the plurality of patches."
… Tu's disclosure is limited to image spatial dimensions, not to time-series patches, feature dimensions of sensor data … Therefore, Tu fails to teach or suggest the feature of claim 1. Maja and Burba do not remedy the deficiencies of Tu.
However, as outlined by the teachings presented in this office action;
Tu teaches applying gated MLP mixing across different directions of patched input, Maja teaches patching input time-series sensor data, Tolstikhin teaches MLP mixing which is channel independent and shares weights across channels, and Huang teaches MLP mixing which mixes across the plurality of patches and across features of the plurality of patches, which, in combination, teaches the entire recited feature of amended claim 1.
Starting on page 7, the applicant further remarks;
Further, amended claim 1 recites "applying a patch-time aggregated hierarchy to guide lowest-level predictions based on aggregated hierarchy signals at a patch-level ... wherein the patch-time aggregated hierarchy comprises generating hierarchical aggregated forecasts from the patch-level predictions generated by the gated MLP mixing and reconciling the generated hierarchical aggregated forecasts."
Burba does not teach or suggest generating hierarchical forecasts from patch-level predictions generated by the gated MLP mixing, performing reconciliation at the patch level, or using hierarchical aggregation signals to guide patch-level predictions. Rather, Burba performs reconciliation after forecasts are generated for hierarchy levels, rather than operating on patch-level predictions generated by an MLP
Therefore, Burba fails to teach or suggest the above-mentioned features of claim 1. Tu and Maja do not remedy the deficiencies of Burba.
However, as outlined by the teachings presented in this office action;
Tu teaches predictions generated by the gated MLP mixing, Zhang teaches applying a patch-time aggregated hierarchy, wherein the patch-time aggregated hierarchy comprises generating hierarchical aggregated forecasts from the patch-level predictions, and Burba teaches applying an aggregated hierarchy to guide lowest-level predictions based on aggregated hierarchy signals at a level, wherein the aggregated hierarchy comprises generating hierarchical aggregated forecasts from predictions at lower levels, and reconciling the generated hierarchical aggregated forecasts, which, in combination, teaches the recited limitation of amended claim 1.
Thus, amended claim 1, and dependent claims 2-4, 6-10, are not patentably distinct over the cited references.
Applicant respectfully submits that the arguments presented for independent claim 1 are equally applicable for independent claims 11 and 20 by virtue of analogous amendments made therein.
As a result, independent claims 1, 11, and 20 are accordingly patentably distinct over the cited references.
Other dependent claims 2-4, 6-10, 12, and 14-19 are patentably distinct for at least the reasons discussed with respect to the respective independent claims 1 and 11. Claims 5 and 13 are canceled herein, thereby rendering their rejection moot.
The arguments of claim 1 are equally applicable to corresponding independent claims 11 and 20, thus, using the same reasoning, amended claims 11 (and dependent claims 12, 14-19) and 20 are not patentably distinct over the cited references.
With the addition of the references Zhang, Huang, and Tolstikhin teaching the subject matter introduced in the amendments, the rejections under 35 U.S.C. 103 stand.
CONCLUSION
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Matthew Alan Cady whose telephone number is (571) 272-7229. The examiner can normally be reached Monday - Friday, 7:30 am - 5:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached on (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC)
at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW ALAN CADY/ Examiner, Art Unit 2145
/CESAR B PAULA/ Supervisory Patent Examiner, Art Unit 2145