DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
This office action is in response to the amendment filed on 06/19/2026. Claims remain pending in the application. Claims 1 and 19-20 are independent.
Specification
Applicant's amendment to specification corrects previous objections; therefore, the previous objections are withdrawn.
Claim Objections
Applicant's amendment to claims corrects previous objections; therefore, the previous objections are withdrawn.
Claim Rejections - 35 USC § 112
Applicant's amendment to claims corrects previous rejections; therefore, the previous rejections are withdrawn.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 10, and 12-20 are rejected under 35 U.S.C. 103 as being unpatentable over ZHANG (CN 109598387 A, pub. date: 04/19/2019), hereinafter ZHANG in view of Chen et al. ("Transformer Encoder With Multi-Modal Multi-Head Attention for Continuous Affect Recognition", IEEE TRANSACTIONS ON MULTIMEDIA, VOL. 23, Nov. 18, 2021, pp. 4171-4183), hereinafter Chen.
Independent Claims 1 and 19-20
ZHANG discloses a computer-implemented method comprising: inputting target series data into a transformer comprising an encoder and a decoder, the transformer having been trained on first and second datasets comprising different first and second modalities, respectively, the second dataset comprising time series data, the encoder comprising separate first and second modality streams for analyzing the first and the second datasets, respectively (ZHANG, ¶¶ [0007]-[0013] with FIG. 1: provide a method and system for stock price prediction using social text and stock price sequence data; provide a usable network framework, namely a method and prediction system for modeling and predicting stock trends using discrete stock price data and text information in social networks; the first step is to select a dataset, crawl stock closing sequence data and corresponding social text datasets such as Twitter, and preprocess the text data; the second step involves using word vectors to transform text sequences into vector feature representations for social texts, and performing three-class classification on continuous stock price sequences to transform them into discrete data representations for stock price sequences; the fourth step is to split the dataset, use the training samples to learn the parameters of the network model, and use the validation set to fine-tune the parameters; ¶¶ [0014]-[0020] with FIG. 2: the stock price information refers to the collected stock market closing price data, and the stock price closing sequence data refers to the stock price data after preprocessing the original stock price information; the preprocessed stock price closing sequence data is used as the input of the model; preferably, this refers to crawling the closing stock price information of the S&P 500 index on Yahoo Finance; social text information refers to information about stocks on various online social platforms, including Twitter, Weibo, WeChat, and others; preferably, it refers to social text information about stocks in the S&P 500; social text information refers to raw text information, while social text data refers to preprocessed text information; in the first step, text data preprocessing refers to the process of removing stop words, special symbols, and links from the crawled social text data (e.g., Twitter text data) because some words or characters contain no information; ¶¶ [0021]-[0027] with FIG.2: in the second step, the text sequence is transformed into a vector feature representation using word vectors, generated according to the following steps: (a) the preprocessed social text data is trained using the word vector model word2vec to learn the word vector representation of each word in the entire text database; the dimension of the word vector is denoted as De; (b) generate vector representations at the social text level; e.g., taking a social media post (Twitter) about a stock as an example, based on the obtained word vectors, average pooling is performed on each dimension of the word vectors of all words in the social media post; the word vector matrix of Nword words in the social text is then subjected to dimensional average pooling to obtain a De dimensional social text representation; (c) generate a vector representation at the day level; e.g., consider social media text related to stocks on a particular day; after obtaining the social text-level vector representations (e.g., Twitter-level representations) of Ntweet social texts (e.g., Twitter) according to the method in step (b) above, for the day level stock text matrix representation with Ntweet*De as the dimension, max, min, and average pooling operations are applied to each dimension of the word vectors to obtain a 3*De stock text representation for that day; realize the transformation of text sequences into vector feature representations using word vectors; in the second step, the operation of performing three-classification processing on the continuous stock price value sequence means that, for the crawled original closing stock price sequence features, if the closing price of the day is higher than the closing price of the previous day, then "+1" is used as the stock price feature of the day; otherwise, "-1" is used as the stock price feature of the day; if the closing price of the day is the same as the closing price of the previous day, then "0" is used as the stock price feature of the day; the continuous stock price sequence features are transformed into a three-class sequence feature taken from {+1, 0, -1}; ¶ [0028] with FIG. 3: use an attention mechanism consisting of an encoder and a decoder: at the encoder end, the attention mechanism is used to select relevant external stocks, and at the decoder end, relevant sequence features are selected for the entire sequence; ¶¶ [0059] and [0066]-[0067] with FIG. 1: in the fourth step, splitting the dataset refers to dividing the entire stock price sequence data set into social text datasets (such as Twitter text datasets) according to time, using the split training set to train the model parameters, and using the validation set to fine-tune the parameters; during model training, the parameters are constrained by using a dropout network and L2 regularization of the parameters to prevent overfitting; in the second step, price data is classified into three categories, text information is pooled using word vectors, and the two parts of data are modeled separately; the key point is to use a long short-term memory network to obtain the hidden state of the memory unit and extract the relationship between stocks and sequences; ¶¶ [0069]-[0072] with FIG. 4: a stock price prediction system utilizing stock price information and social text information, which comprises input the representation unit to preprocess the original stock closing data and Twitter text data respectively, discretize the original stock closing data, and use word vectors to serialize the Twitter text data; ¶¶ [0080]-[0085] with FIG. 1: provide a method for prediction using social text and stock price sequence data; the first step is to select the corresponding Twitter social text dataset for the stock price closing sequence dataset required for the task, and perform preprocessing such as noise removal on the text data; the second step involves using word vectors to transform the preprocessed text sequence into vector feature representations, and then classifying the stock price sequence data into discrete data representations; the fourth step is to split the dataset, train the dataset, and fine-tune the parameters; ¶¶ [0086]-[0089] with FIG. 4: a stock price prediction system comprising: input representation unit where (a) the original stock closing price data and Twitter text data are preprocessed separately; (b) the original stock closing price data is discretely classified into three categories; and (c) the vectorized representation of the Twitter text data is generated using word vector; ¶¶ [0091]-[0100] with FIG. 2: scrape the stocks in the S&P 500 index from Yahoo Finance and extract the closing price of each stock each day; using the stock tag "$" as the crawling keyword, the Python framework tool tweepy was used to crawl Twitter text related to S&P 500 stocks; filter special characters from the crawled Twitter text, remove stop words that have no information content, and remove URL information that appears in a large amount of text; discretization of stock price series: (a) for the crawled stock price sequence data, if the closing price of the day is higher than the closing price of the previous day, "+1" is used as the stock price feature of the day; otherwise, "-1" is used; if the closing price of the day is the same as the closing price of the previous day, 0 is used as the stock price feature of the day; and (b) the continuous stock price sequence features are transformed into a three-class sequence feature derived from {+1, 0, -1}; vectorized representation of text: (a) first, for the denoised Twitter text, the word vector representation of each word in the text library is learned using the word vector model word2vec; (b) taking a stock tweet as an example, based on the obtained word vectors, perform average pooling on the word vectors of all words in each dimension; and (c) generate a vector representation at the day level; this vector representation serves as the text input representation of the model; ¶¶ [0122] and [0127] with FIGS. 3-4: the entire dataset is split according to the timeline, with the training set, validation set, and test set in a ratio of 8: 1: 1; the training set is used to train and learn the parameters of the entire model, and the validation set is used to fine-tune the model parameters; during model training, to prevent overfitting, a dropout network and L2 regularization of the parameters are used to limit the training size of the parameters),
(ZHANG, ¶¶ [0007]-[0013] with FIG. 1: jointly model stock prices and text, and to bidirectionally calculate cross-modal attention weights; employs a cross-modal bidirectional attention mechanism to model stock price data and social text, effectively extracting important sequence information; the third step involves modeling the stock price sequence data and social text datasets such as Twitter using recurrent neural networks; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, allowing them to learn to extract stock price sequences and social text sequences relevant to the prediction target; ¶¶ [0028]-[0058] and [0068] with FIG. 1 and 3: in the third step, recurrent neural networks are used to model the stock price sequence data and the social text dataset, respectively; among them, the modeling of stock price sequence data is as follows: a recurrent neural network is modeled using external stock prices and target historical price sequence data; the core of the model is to use an attention mechanism consisting of an encoder and a decoder: at the encoder end, the attention mechanism is used to select relevant external stocks, and at the decoder end, relevant sequence features are selected for the entire sequence; use an attention mechanism at the encoding end to select relevant external stocks: (a) input a sequence of M external stocks of length T, [X1, …, XM], where each stock is a vector representation of length T; (b) calculate attention weights using the input; (c) attention weights are used to select external stock price features relevant to the prediction of stocks; this feature is used to update the state value of the memory cell/unit; at the decoding end, relevant sequence features are selected for the entire sequence: (a) the state features of the memory unit at each time step input from the encoding end and the input sequence features of the text module are used to select the sequence features related to the predicted value from the entire sequence using an attention mechanism; wherein, the attention weight is calculated; (b) calculate the weighted sum representation of the state sequence using attention weights; The weighted sum of the state sequence is used to update the memory cell/unit state at the decoding end together with the historical time series of the target stock; in the third step, modeling the social text (Twitter text) dataset using a recurrent neural network refers to: modeling the vectorized social text sequence representation obtained from the preprocessing in the first step using a long short-term memory network; the input is [E1, …, ET], which represents a sequence of T-length Twitter text vectors of the target stock, i.e., the vector representation of social text (e.g., Twitter text) after preprocessing according to the method in the second step; the weighted sequence sum representation Cd in the stock price sequence module is used to participate in the calculation of text attention weights; this text attention weight can be calculated as a weighted sum of features of the text sequence; Feature Ctext is used to update the state of memory cells/units; in the third step, the use of a bidirectional cross-modal attention mechanism means that, at the decoding end of the stock price network module, the input features of the text [E1, …, ET] are used to help train the sequence attention weights; in the text network module, the weighted sum representation of the hidden states Cd calculated in the stock price module is used to update the text attention weights; therefore, at each moment in the sequence, both modules bidirectionally compute their respective attention weights using cross-modal data from each other; in the third step, the bidirectional cross-modal attention mechanism integrates stock price and text data; its core is the information interaction between the stock price module decoding end and the text module; the method adopted is to utilize the hidden state sequence of the decoding end and the input sequence of the text respectively; ¶¶ [0069]-[0072] with FIG. 4: a stock price prediction system utilizing stock price information and social text information, which comprises text and price sequence modeling unit to perform sequence modeling on the stock price data and text data of the input representation, calculate the attention weight of the two data parts using mutual information, and select relevant input representations; ¶ [0083] with FIG. 1: the third step involves using recurrent neural networks to model the stock price sequence data and the Twitter text dataset, respectively; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, extracting sequence features relevant to the prediction target; ¶¶ [0086]-[0089] with FIG. 4: a stock price prediction system comprising: text and price series modeling unit, where (a) a long short-term memory network is used to perform sequence modeling on the input representations of stock price data and text data; and (b) a bidirectional attention mechanism is used to select relevant input representations by calculating the attention weights of the two data parts using mutual information; ¶¶ [0101]-[0121] with FIGS. 1 and 3: use the LSTM module in TensorFlow as a basis to perform sequence modeling on the two parts of the sequence data; for stock price sequence data modeling, modeling a recurrent neural network using external stock price sequences and historical price sequences of the target stock; at the encoding end, (a) the input sequence consists of M external stocks of length T, [X1, …, XM], where each stock is a vector representation of length T; (b) calculate attention weights using the input; (c) attention weights are used to select external stock price features relevant to the prediction of stocks; at the decoding end, the state features of the memory unit/cell at each time step input from the input encoder and the input sequence features of the text module are used to select the sequence features in the entire sequence that are relevant to the predicted value using an attention mechanism; attention weights are used to select the relevant memory cell states at the encoder end; for Text sequence data modeling, after obtaining a vectorized text sequence representation through preprocessing, a Long Short-Term Memory (LSTM) network is used to model the text sequence; the initial input is [E1, …, ET], representing a sequence of T-length Twitter text vectors of the target stock, which is the vector representation of the preprocessed Twitter text; the text input features and the weighted sequence representation Cd from the stock price sequence module are used to calculate the text attention weights) (NOTE: ZHANG only teaches (a) the second modality stream performing feature-level attention, intra-modal multi-head attention, and inter-modal multi-head attention; and (b) the first modality stream performing inter-modal multi-head attention); and
in response to the inputting, receiving from the transformer an inferred variable related to the target series data (ZHANG, ¶ [0013] with FIG. 1: the fifth step is to use a network model based on bidirectional cross-modal attention to predict stock price trends in the target data; ¶¶ [0060]-[0065] with FIG. 1: in the fifth step, a network model based on bidirectional cross-modal attention is used to predict the target stock price trend; following the method in step three, obtain the stock price sequence and social text sequence related to the prediction target; i.e., obtain the hidden unit state of the stock price sequence module decoder and the text module at each moment [
h
1
d
e
, …,
h
T
d
e
], [
h
1
t
e
x
t
, …,
h
T
t
e
x
t
]; take the state representation of the last day from the two features in step (1) above and concatenate them to obtain [
h
T
d
e
,
h
T
t
e
x
t
]; prediction is performed using the features of splicing, as follows:
Y
=
ο
(
v
o
W
o
h
T
d
e
;
h
T
t
e
x
t
+
b
o
+
b
v
)
, where sigmoid is used as the activation function σ;
v
o
,
W
o
,
b
o
, and
b
v
are the parameters that need to be trained in the network; ¶¶ [0069]-[0074] with FIG. 4: a stock price prediction system utilizing stock price information and social text information, which comprises prediction generation unit to obtain the hidden state of the last day in the partial text and price sequence modeling from the text and price sequence modeling unit and splice it, connect it to the two-layer fully connected layer and finally outputs the sigmoid activation; integrate the stock price module and the social text module to calculate their respective attention weights at each time step of the sequence, requiring the attention calculation sequence to be homogeneous; utilize social text sequence data, combined with the stock price information of the target stock and external stocks, to jointly predict the target stock trend, and based on the two-way attention interaction, it can select important stock price sequence and text sequence features respectively; ¶ [0085] with FIG. 1: the fifth step is to predict the stock price trend of the target data based on a network model with bidirectional cross-modal attention; ¶¶ [0086]-[0089] with FIG. 4: a stock price prediction system comprising: prediction generation unit to obtain the hidden state of the long short-term memory network of the last day in the part of the text and price sequence modeling from text and price series modeling unit. and splice it, then connect it to the output of the sigmoid activation; ¶¶ [0123]-[0126] with FIGS. 3-4: during the prediction process, the hidden unit states [
h
1
d
e
, …,
h
T
d
e
], [
h
1
t
e
x
t
, …,
h
T
t
e
x
t
] at each time step of the stock price sequence module and the text module are used to extract the state features of the last time step and concatenate them into [
h
T
d
e
,
h
T
t
e
x
t
] which is then used for prediction:
Y
=
ο
(
v
o
W
o
h
T
d
e
;
h
T
t
e
x
t
+
b
o
+
b
v
)
, where sigmoid is used as the activation function σ;
v
o
,
W
o
,
b
o
, and
b
v
are the parameters that need to be trained in the network).
ZHANG further discloses a computer program product comprising: a computer readable storage medium (excluding transitory signal per se. in ¶ [0020] of the specification) having program code embodied therewith, the program code executable by a processor to perform the method described above (ZHANG, ¶¶ [0007]-[0013], [0069]-[0072], and [0086]-[0089] with FIG. 4: inherited in the system used to train the stock price sequence data and social text datasets, and predict stock price trends in the target data).
ZHANG further discloses a computer system comprising: a processor operatively coupled to memory, and an artificial intelligence (Al) platform operatively coupled to the processor, the Al platform comprising a transformer and one or more tools configured to interface with the transformer, including the method described above (ZHANG, ¶¶ [0007]-[0013], [0069]-[0072], and [0086]-[0089] with FIG. 4: inherited in the system used to train the stock price sequence data and social text datasets, and predict stock price trends in the target data).
ZHANG fails to explicitly disclose each of the first and second modality streams respectively performing feature-level attention, intra-modal multi-head attention, and inter-modal multi-head attention.
Chen teaches a system and method using attention mechanism (Chen, Abstract in Page 4171), wherein each of the first and second modality streams respectively performing feature-level attention, intra-modal multi-head attention, and inter-modal multi-head attention (Chen, Abstract in Page 4171: first introduce the transformer-encoder with a self-attention mechanism and propose a Convolutional Neural Network-Transformer Encoder (CNNTE) framework to model the temporal dependency for single modal affect recognition; further, to effectively consider the complementarity and redundancy between multiple streams, propose a Transformer Encoder with Multi-modal Multi-head Attention (TEMMA) for multi-modal affect recognition; TEMMA allows to progressively and simultaneously refine the inter-modality interactions and intra-modality temporal dependency; the learned multi-modal representations are fed to an Inference Sub-network with fully connected layers to estimate the affective state; Section I in Pages 4171-4172: the transformer, which relies entirely on a self-attention mechanism, demonstrated its effectiveness to draw global dependencies over a sequence; multi-modal transformer frameworks have been built upon the transformer-encoder with a pair-wise cross-modality attention module performing on the low-level descriptors or on the high-level representations; these frameworks considered temporal dynamics and multi-modal dynamic interactions separately; to explore the dynamic interactions between different modalities along with the temporal dependency, an intermediate level multi-modal fusion based on the transformer is proposed; the model constructs multiple feature streams, each consisting of factorized multi-modal features; then the multi-head attention is performed on each feature stream to learn multi-modal dynamic interactions along with temporal dependency; based on the transformer], address the dynamic temporal dependency and multi-modal fusion challenges of continuous affect recognition, as follows: (a) to handle the long-range dependencies, propose to predict affective states by combing one dimensional CNN with the Transformer-Encoder (CNN-TE), wherein the CNN is used to locally aggregate context information, while the transformer-encoder models the long-range dependencies with the attention mechanism; (b) to deal with the dynamic importance of multi-modal feature streams, propose the Multi-modal Multi-head Attention (MMA), which can be easily inserted into the transformer-encoder to promote the dynamic interactions between different modalities and learn their correlations; (c) furthermore, suggest a Transformer-Encoder with Multi-modal Multi-head Attention (TEMMA) and propose a CNN-TEMMA framework, to progressively and simultaneously refine inter-modality interactions and intramodality dependencies, wherein the CNN sub-network extracts and aggregates local temporal features via causal convolution, and the TEMMA sub-network models multi-modal interactions and temporal dynamics, with spatial (audio and visual features) and temporal attention, to learn high-level representations from each modality; and (d) finally, an inference sub-network, with fully connected layers, is used to predict the affective state from the learned multi-modal representations; Section II.B in Page 4173: [55] propose Tensor Fusion Network learn intra-modality and inter-modality dynamics in an end-to-end fashion, wherein intra-modality dynamics is obtained via Modality Embedding Subnetworks, while inter-modality dynamics is obtained via the 3-fold Cartesian product between each modality embeddings to capture bi-modal and tri-modal interactions; [57] propose the Multimodal Transformer Networks (MTN) to generate conversational responses to queries of humans for a video-grounded dialogue system (VGDS), wherein MTN comprises three major components: (a) transformer-encoder layers to map each sequence of tokens (text, video) into a sequence of continuous representation; (b) transformer-decoder layers to perform reasoning over multiple encoded features through a multi-head attention mechanism; and finally, (c) auto-encoder layers that are used to focus on query-related video features in an unsupervised manner; the authors of [38] extended the transformer of with a pairwise cross-modal attention module to learn representations from different paired multi-modal features; the representations from each pair are further fed into different transformer-encoders to achieve temporal dependency; the authors of [39] proposed to firstly perform the transformer-encoder on each modality individually, then apply the pairwise multi-modal interaction on the high-level multi-modal representations; the authors of [40] propose an intermediate level multi-modal fusion based on the transformer-encoder; concretely, uni-modal, bi-modal and tri-modal feature sequences are formed as one feature sequence and reconstructed into multiple streams with a factorized multi-modal features; then, a multi-head attention is performed on each factorized stream individually, and the output streams are added into one feature sequence in a point-wise manner; we propose an intermediate level multi-modal fusion scheme built upon the standard Transformer network to efficiently learn intra-modality dynamics and inter-modality dynamics; a novel Multi-modal Multi-head Attention (MMA) module, which learns relationships between several multi-modal sequences, is first presented; then the MMA module is followed by a temporal multi-head attention (TMA) module for each feature stream, composing a CNN and transformer-encoder with multi-modal multi-head attention (CNN-TEMMA) framework; this framework progressively learns multi-modal interactions and temporal dependency; the proposed CNN-TEMMA framework can be used not only for audiovisual continuous affect recognition, but also for other multi-modal feature learning tasks with dynamic sequences as inputs; Section II.C in Pages 4173-4174: for enforcing attention on important cues within a sequence, attention mechanisms have been integrated into sequence-to-sequence models; Vaswani et al. [36] proposed a self-attention based sequence to sequence model, namely the transformer which employs multi-head attention to calculate correlations over a temporal sequence; employ the transformer-encoder to model the temporal dynamics as it produces an attentive feature sequence with the same length as the input sequence, being vital for continuous affect recognition; Section III with FIGS. 1-2 in Pages 4174-4175: introduce our proposed uni-modal affect recognition framework, which combines a CNN and a Transformer-Encoder (CNN-TE), as illustrated in Fig. 1; the extracted feature vectors are fed into the 1-dimensional temporal convolutional network for aggregating local temporal context information; the output of the 1-D CNN is input into the transformer-encoder to achieve long-range dependencies with dynamic attention weights; finally, an inference sub-network, with fully connected layers, is used to estimate the valence or arousal dimension of the affective state; a 1-D temporal convolution network is adopted to encode the temporal information from the input feature sequence; in the model, use causal convolutions to respect the temporal order during the learning process, wherein "causal" means that the activations computed for a particular time step do not depend on the activations from the future time steps; since the transformer-encoder contains neither recurrence nor convolution, to make use of the time step order of the feature vectors within the sequence, add the position information to the output of the 1-D temporal convolutional network; as shown in Fig. 1, the transformer-encoder is composed of N identical blocks, each consisting of a multi-head attention module followed by a fully connected feed-forward module, which employs multi-head attention (MA) to calculate the temporal dependency (therefore we name it as temporal MA (TMA)); as introduced in [36], the attention function can be described as performing a query on a set of key-values pairs, to generate an output; the output is computed as a weighted sum of values, where the weights is calculated by a compatibility function of the query with the corresponding key; Fig. 2 illustrates the multi-head attention module used in the transformer-encoder, which employs the scaled dot-product attention on the queries (Q), keys (K), and values (V ) under each head; in practice, Q is a set of queries of the whole sequence, packed together in a matrix; the keys and values are also packed together into matrices K and V; Q, K, V are first projected h times in different subspace headi (i ∈ [1, h]) with different learned linear projections
W
i
Q
,
W
i
K
,
W
i
V
of dimensions dk, dk, dv, respectively; then the scaled dot-product attention is performed in parallel on headi of queries, keys, and values; concretely, the scaled dot-product attention computes the dot products (MatMul) of the scaled query and keys; the result is multiplied by a pre-defined attention weights mask, before applying the softmax function to obtain the weights A on the values; such attention block can be mathematically described as eqn. (1), where ∗ represents the Hadamard product, and Mask refers to a bidirectional attention weights mask, which can be regarded as a hard attention mask making the encoder ignore the information from far history and consider the near future; the multi-head attention module (TMA) is given as eqn. (2); apart from the multi-head attention module and the fully connected feed-forward module, the transformer-encoder contains a residual connection followed by a normalization layer (LN); as illustrated in Fig. 1, the transformer-encoder maps the input sequence of low-level descriptors, to an output sequence of high-level representations; the learned high-level representations are fed to an Inference sub-network to estimate the arousal or valence dimension of the affective state, which is composed of two fully connected layers with a non-linear activation layer in between; the number of nodes of the last fully connected layer is half of the dimension of the high-level representations; Section IV in Pages 4175-4176 with FIG. 4 in Page 4176 and FIG. 5 in Page 4171: the proposed CNN and transformer-encoder with multi-modal multi-head attention (CNN-TEMMA) framework consists of three modules: a) the Input Embedding sub-network with 1D convolution, which outputs the embedded feature sequence, to which we inject the position information; b) the Multi-modal Encoder sub-network with stacked N identical encoder blocks, wherein for each block, insert the proposed multi-modal multi-head attention (MMA) module to model the inter-modality interactions, followed by the temporal multi-head attention (TMA) module to model the intra-modality dependencies; c) the Inference sub-network, in which the encoded high-level representations from the different modalities are concatenated into one feature vector linked to a fully connected deep neural network for affective state estimation; for each modality, the input embedding sub-network is the same as the one described in Section III-A for the uni-modal affect recognition; it consists of two parts: local context information aggregation via a 1D temporal convolutional network, and the position information embedding; the Multi-modal Encoder sub-network of Fig. 4, is composed of N stacked encoder blocks, wherein the encoder block comprises a Multi-modal Multi-head Attention (MMA) module followed by the original Multi-head Attention (TMA) of Fig. 1; such architecture not only dynamically enriches each modality with complementary information from the other modalities but also directs TMA on searching Intra-modality dynamics; furthermore, by stacking the encoder blocks, the encoder could progressively refine the inter-modality interactions and intra-modality dynamics to learn high-level semantic representation with temporal contextual information as well as complementary multi-modal information; Temporal Multi-Head Attention (TMA) Module is performed individually on each modality to capture the temporal dependency by following the same principles as described in Section III-B; to obtain complementary information from different modalities, propose a Multi-Modal Multi-Head Attention (MMA) module to compute inter-modality interactions, as illustrated in Fig. 5, which the same as that in the TMA module described in Section IV-B1 following the self-attention paradigm; the MMA will calculate the attention weights indicating the inter-modality correlations, which composes a matrix of weights in
R
m
×
m
at all time steps t ∈ [1, n]; the structure of the inference sub-network for multi-modal affect recognition is the same as that for the single modal model; the only difference is the input feature vector, wherein the outputs of the multi-modal encoder sub-network, i.e., a set of high-level representations from multiple modalities, are concatenated to form an input feature vector to the inference sub-network to estimate the value of arousal or valence).
ZHANG and Chen are analogous art because they are from the same field of endeavor, a system and method using attention mechanism. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of Chen to ZHANG. Motivation for doing so would extend the framework for any multi-modal feature learning tasks with dynamic sequences as inputs ( and provide better performance1.
Claim 2
ZHANG in view of Chen discloses all the elements as stated in Claim 1 and further discloses wherein one or more identified first and second temporal features identified by the intra-modal multi-head attentions of the first and second modality streams are input into the respective inter-modal multi-head attention to identify one or more cross modality relationships (ZHANG, ¶¶ [0007]-[0013] with FIGS. 1 and 3: jointly model stock prices and text, and to bidirectionally calculate cross-modal attention weights; employ a cross-modal bidirectional attention mechanism to model stock price data and social text, effectively extracting important sequence information; the stock price prediction method based on a bidirectional cross-modal attention network model; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, allowing them to learn to extract stock price sequences and social text sequences relevant to the prediction target; use a network model based on bidirectional cross-modal attention to predict stock price trends in the target data; ¶ [0039] with FIG. 3: the state features of the memory unit at each time step input from the encoding end and the input sequence features of the text module are used to select the sequence features related to the predicted value from the entire sequence using an attention mechanism; ¶¶ [0058], [0060], and [0068] with FIGS.1 and 3: the use of a bidirectional cross-modal attention mechanism means that, at the decoding end of the stock price network module, the input features of the text [E1, …, ET] are used to help train the sequence attention weights; in the text network module, the weighted sum representation of the hidden states Cd calculated in the stock price module is used to update the text attention weights; therefore, at each moment in the sequence, both modules bidirectionally compute their respective attention weights using cross-modal data from each other; a network model based on bidirectional cross-modal attention is used to predict the target stock price trend; the bidirectional cross-modal attention mechanism integrates stock price and text data; its core is the information interaction between the stock price module decoding end and the text module; the method adopted is to utilize the hidden state sequence of the decoding end and the input sequence of the text respectively; ¶¶ [0083] and [0085] with FIG. 3: involve using recurrent neural networks to model the stock price sequence data and the Twitter text dataset, respectively; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, extracting sequence features relevant to the prediction target; predict the stock price trend of the target data based on a network model with bidirectional cross-modal attention) (Chen, Section IV in Pages 4175-4176 with FIG. 4 in Page 4176 and FIG. 5 in Page 4171: for a given Modalityj, denote by
Q
j
t
the feature vector at time t, and Qj = (
Q
j
1
, …
Q
j
t
, …,
Q
j
n
) the full feature sequence, then Kj = Vj = Qj are the input feature sequences to the TMA module of modality j (Fig. 5); for each modality j, the process can be described as eqn. (3), wherein i ∈ [1,h] denotes the ith head of the total number of h heads, j ∈ [1,m] denotes the jth modality, t ∈ [1,n] denotes the frame, and n the number of frames; the projections are parameter matrices:
W
j
,
i
T
Q
∈
R
d
m
o
d
e
l
×
d
k
,
W
j
,
i
T
K
∈
R
d
m
o
d
e
l
×
d
k
,
W
j
,
i
T
V
∈
R
d
m
o
d
e
l
×
d
v
and
W
j
,
i
T
O
∈
R
h
d
v
×
d
m
o
d
e
l
; the superscripts (T) denotes matrices of the temporal multi-head attention module; propose a MMA module to compute inter-modality interactions, as illustrated in Fig. 5; Let
Q
j
t
the feature vector of modality j at time step t, and Qt = (
Q
1
t
, …
Q
j
t
, …,
Q
m
t
) the multi-modal feature sequence at time step t, then Kt = Vt = Qt are the input multi-modal feature sequences to the MMA module at time step t (Fig. 5), which is the same as that in the TMA module (section IV-B1) following the self-attention paradigm, here the input feature sequences are organized in a multi-modal manner under each time step; the MMA will calculate the attention weights indicating the inter-modality correlations, which composes a matrix of weights in
R
m
×
m
at all time steps t ∈ [1,n]; in practice, for each modality, the same as for TMA, the input is projected into multiple subspaces but with different parameters W(M)Q, W(M)K, W(M)V, wherein the superscript (M) denotes matrices of the MMA module; then, at each time step t, construct the multi-modal queries, keys and values (
Q
i
t
,
K
i
t
,
V
i
t
) under subspace headi followed by the scaled dot-product attention; finally, the output values under each subspace headi are concatenated and linearly projected, resulting in the final values as shown in eqns. (4)-(6)).
Claim 3
ZHANG in view of Chen discloses all the elements as stated in Claim 2 and further discloses wherein the identified one or more cross modality relationships and signals from the feature-level attentions are combined to produce the output of the encoder (Chen, Section IV.C in Page 4176: the structure of the inference sub-network for multi-modal affect recognition is the same as that for the single modal model; the only difference is the input feature vector, wherein the outputs of the multi-modal encoder sub-network, i.e., a set of high-level representations from multiple modalities, are concatenated to form an input feature vector to the inference sub-network to estimate the value of arousal or valence) (ZHANG, ¶¶ [0007]-[0013] with FIG. 1: jointly model stock prices and text, and to bidirectionally calculate cross-modal attention weights; employs a cross-modal bidirectional attention mechanism to model stock price data and social text, effectively extracting important sequence information; the third step involves modeling the stock price sequence data and social text datasets such as Twitter using recurrent neural networks; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, allowing them to learn to extract stock price sequences and social text sequences relevant to the prediction target; ¶¶ [0028]-[0058] and [0068] with FIG. 1 and 3: in the third step, recurrent neural networks are used to model the stock price sequence data and the social text dataset, respectively; among them, the modeling of stock price sequence data is as follows: a recurrent neural network is modeled using external stock prices and target historical price sequence data; the core of the model is to use an attention mechanism consisting of an encoder and a decoder: at the encoder end, the attention mechanism is used to select relevant external stocks, and at the decoder end, relevant sequence features are selected for the entire sequence; use an attention mechanism at the encoding end to select relevant external stocks: (a) input a sequence of M external stocks of length T, [X1, …, XM], where each stock is a vector representation of length T; (b) calculate attention weights using the input; (c) attention weights are used to select external stock price features relevant to the prediction of stocks; this feature is used to update the state value of the memory cell/unit; at the decoding end, relevant sequence features are selected for the entire sequence: (a) the state features of the memory unit at each time step input from the encoding end and the input sequence features of the text module are used to select the sequence features related to the predicted value from the entire sequence using an attention mechanism; wherein, the attention weight is calculated; (b) calculate the weighted sum representation of the state sequence using attention weights; The weighted sum of the state sequence is used to update the memory cell/unit state at the decoding end together with the historical time series of the target stock; in the third step, modeling the social text (Twitter text) dataset using a recurrent neural network refers to: modeling the vectorized social text sequence representation obtained from the preprocessing in the first step using a long short-term memory network; the input is [E1, …, ET], which represents a sequence of T-length Twitter text vectors of the target stock, i.e., the vector representation of social text (e.g., Twitter text) after preprocessing according to the method in the second step; the weighted sequence sum representation Cd in the stock price sequence module is used to participate in the calculation of text attention weights; this text attention weight can be calculated as a weighted sum of features of the text sequence; Feature Ctext is used to update the state of memory cells/units; in the third step, the use of a bidirectional cross-modal attention mechanism means that, at the decoding end of the stock price network module, the input features of the text [E1, …, ET] are used to help train the sequence attention weights; in the text network module, the weighted sum representation of the hidden states Cd calculated in the stock price module is used to update the text attention weights; therefore, at each moment in the sequence, both modules bidirectionally compute their respective attention weights using cross-modal data from each other; in the third step, the bidirectional cross-modal attention mechanism integrates stock price and text data; its core is the information interaction between the stock price module decoding end and the text module; the method adopted is to utilize the hidden state sequence of the decoding end and the input sequence of the text respectively; ¶¶ [0069]-[0072] with FIG. 4: a stock price prediction system utilizing stock price information and social text information, which comprises text and price sequence modeling unit to perform sequence modeling on the stock price data and text data of the input representation, calculate the attention weight of the two data parts using mutual information, and select relevant input representations; ¶ [0083] with FIG. 1: the third step involves using recurrent neural networks to model the stock price sequence data and the Twitter text dataset, respectively; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, extracting sequence features relevant to the prediction target; ¶¶ [0086]-[0089] with FIG. 4: a stock price prediction system comprising: text and price series modeling unit, where (a) a long short-term memory network is used to perform sequence modeling on the input representations of stock price data and text data; and (b) a bidirectional attention mechanism is used to select relevant input representations by calculating the attention weights of the two data parts using mutual information; ¶¶ [0101]-[0121] with FIGS. 1 and 3: use the LSTM module in TensorFlow as a basis to perform sequence modeling on the two parts of the sequence data; for stock price sequence data modeling, modeling a recurrent neural network using external stock price sequences and historical price sequences of the target stock; at the encoding end, (a) the input sequence consists of M external stocks of length T, [X1, …, XM], where each stock is a vector representation of length T; (b) calculate attention weights using the input; (c) attention weights are used to select external stock price features relevant to the prediction of stocks; at the decoding end, the state features of the memory unit/cell at each time step input from the input encoder and the input sequence features of the text module are used to select the sequence features in the entire sequence that are relevant to the predicted value using an attention mechanism; attention weights are used to select the relevant memory cell states at the encoder end; for Text sequence data modeling, after obtaining a vectorized text sequence representation through preprocessing, a Long Short-Term Memory (LSTM) network is used to model the text sequence; the initial input is [E1, …, ET], representing a sequence of T-length Twitter text vectors of the target stock, which is the vector representation of the preprocessed Twitter text; the text input features and the weighted sequence representation Cd from the stock price sequence module are used to calculate the text attention weights).
Claim 10
ZHANG in view of Chen discloses all the elements as stated in Claim 1 and further discloses wherein the first dataset comprises time-stamped textual data and the second dataset comprises numerical time series data (ZHANG, ¶¶ [0009]-[0010] and [0014]-[0020]: select a dataset, crawl stock closing sequence data and corresponding social text datasets such as Twitter, and preprocess the text data; involves using word vectors to transform text sequences into vector feature representations for social texts, and performing three-class classification on continuous stock price sequences to transform them into discrete data representations for stock price sequences; the stock price information refers to the collected stock market closing price data, and the stock price closing sequence data refers to the stock price data after preprocessing the original stock price information; the preprocessed stock price closing sequence data is used as the input of the model; preferably, this refers to crawling the closing stock price information of the S&P 500 index on Yahoo Finance; social text information refers to information about stocks on various online social platforms, including Twitter, Weibo, WeChat, and others; preferably, it refers to social text information about stocks in the S&P 500).
Claim 12
ZHANG in view of Chen discloses all the elements as stated in Claim 10 (see also 112(b) rejection to Claim 12) and further discloses wherein the time-stamped textual data is produced via performing natural language processing on text articles (ZHANG, ¶¶ [0022] and [0098]: the preprocessed social text data is trained using the word vector model word2vec to learn the word vector representation of each word in the entire text database; for the denoised Twitter text, the word vector representation of each word in the text library is learned using the word vector model word2vec) (Chen, Section II.C in Pages 4173-4174: the authors of [59] proposed a new language representation model, the Bidirectional Encoder Representations from Transformers (BERT) which employs the transformer-encoder of [36] to extract contextual information within a text sequence; BERT has been widely adopted and achieved great success in natural language processing tasks).
Claim 13
ZHANG in view of Chen discloses all the elements as stated in Claim 1 and further discloses wherein the intra-modal multi-head attentions extract temporal dependencies between different time steps in a single modality (Chen, Abstract in Page 4171: first introduce the transformer-encoder with a self-attention mechanism and propose a Convolutional Neural Network-Transformer Encoder (CNNTE) framework to model the temporal dependency for single modal affect recognition; further, to effectively consider the complementarity and redundancy between multiple streams, propose a Transformer Encoder with Multi-modal Multi-head Attention (TEMMA) for multi-modal affect recognition; TEMMA allows to progressively and simultaneously refine the inter-modality interactions and intra-modality temporal dependency; the learned multi-modal representations are fed to an Inference Sub-network with fully connected layers to estimate the affective state; Section I in Pages 4171-4172: the transformer, which relies entirely on a self-attention mechanism, demonstrated its effectiveness to draw global dependencies over a sequence; multi-modal transformer frameworks have been built upon the transformer-encoder with a pair-wise cross-modality attention module performing on the low-level descriptors or on the high-level representations; these frameworks considered temporal dynamics and multi-modal dynamic interactions separately; to explore the dynamic interactions between different modalities along with the temporal dependency, an intermediate level multi-modal fusion based on the transformer is proposed; the model constructs multiple feature streams, each consisting of factorized multi-modal features; then the multi-head attention is performed on each feature stream to learn multi-modal dynamic interactions along with temporal dependency; based on the transformer, address the dynamic temporal dependency and multi-modal fusion challenges of continuous affect recognition, as follows: (a) to handle the long-range dependencies, propose to predict affective states by combing one dimensional CNN with the Transformer-Encoder (CNN-TE), wherein the CNN is used to locally aggregate context information, while the transformer-encoder models the long-range dependencies with the attention mechanism; (b) to deal with the dynamic importance of multi-modal feature streams, propose the Multi-modal Multi-head Attention (MMA), which can be easily inserted into the transformer-encoder to promote the dynamic interactions between different modalities and learn their correlations; (c) furthermore, suggest a Transformer-Encoder with Multi-modal Multi-head Attention (TEMMA) and propose a CNN-TEMMA framework, to progressively and simultaneously refine inter-modality interactions and intramodality dependencies, wherein the CNN sub-network extracts and aggregates local temporal features via causal convolution, and the TEMMA sub-network models multi-modal interactions and temporal dynamics, with spatial (audio and visual features) and temporal attention, to learn high-level representations from each modality; and (d) finally, an inference sub-network, with fully connected layers, is used to predict the affective state from the learned multi-modal representations; Section II.B in Page 4173: the authors of [38] extended the transformer of with a pairwise cross-modal attention module to learn representations from different paired multi-modal features; the representations from each pair are further fed into different transformer-encoders to achieve temporal dependency; propose an intermediate level multi-modal fusion scheme built upon the standard Transformer network to efficiently learn intra-modality dynamics and inter-modality dynamics; a novel Multi-modal Multi-head Attention (MMA) module, which learns relationships between several multi-modal sequences, is first presented; then the MMA module is followed by a temporal multi-head attention (TMA) module for each feature stream, composing a CNN and transformer-encoder with multi-modal multi-head attention (CNN-TEMMA) framework; this framework progressively learns multi-modal interactions and temporal dependency; the proposed CNN-TEMMA framework can be used not only for audiovisual continuous affect recognition, but also for other multi-modal feature learning tasks with dynamic sequences as inputs; Section III.B with FIGS. 1-2 in Pages 4174-4175: as shown in Fig. 1, the transformer-encoder is composed of N identical blocks, each consisting of a multi-head attention module followed by a fully connected feed-forward module, which employs multi-head attention (MA) to calculate the temporal dependency (therefore we name it as temporal MA (TMA)); as introduced in [36], the attention function can be described as performing a query on a set of key-values pairs, to generate an output; the output is computed as a weighted sum of values, where the weights is calculated by a compatibility function of the query with the corresponding key; Fig. 2 illustrates the multi-head attention module used in the transformer-encoder, which employs the scaled dot-product attention on the queries (Q), keys (K), and values (V ) under each head; in practice, Q is a set of queries of the whole sequence, packed together in a matrix; the keys and values are also packed together into matrices K and V; Q, K, V are first projected h times in different subspace headi (i ∈ [1, h]) with different learned linear projections
W
i
Q
,
W
i
K
,
W
i
V
of dimensions dk, dk, dv, respectively; then the scaled dot-product attention is performed in parallel on headi of queries, keys, and values; concretely, the scaled dot-product attention computes the dot products (MatMul) of the scaled query and keys; the result is multiplied by a pre-defined attention weights mask, before applying the softmax function to obtain the weights A on the values; such attention block can be mathematically described as eqn. (1), where ∗ represents the Hadamard product, and Mask refers to a bidirectional attention weights mask, which can be regarded as a hard attention mask making the encoder ignore the information from far history and consider the near future; the multi-head attention module (TMA) is given as eqn. (2); apart from the multi-head attention module and the fully connected feed-forward module, the transformer-encoder contains a residual connection followed by a normalization layer (LN); as illustrated in Fig. 1, the transformer-encoder maps the input sequence of low-level descriptors, to an output sequence of high-level representations; Section IV.B in Pages 4175-4176 with FIG. 4 in Page 4176 and FIG. 5 in Page 4171: the Multi-modal Encoder sub-network of Fig. 4, is composed of N stacked encoder blocks, wherein the encoder block comprises a Multi-modal Multi-head Attention (MMA) module followed by the original Multi-head Attention (TMA) of Fig. 1; such architecture not only dynamically enriches each modality with complementary information from the other modalities but also directs TMA on searching Intra-modality dynamics; furthermore, by stacking the encoder blocks, the encoder could progressively refine the inter-modality interactions and intra-modality dynamics to learn high-level semantic representation with temporal contextual information as well as complementary multi-modal information; Temporal Multi-Head Attention (TMA) Module is performed individually on each modality to capture the temporal dependency by following the same principles as described in Section III-B; to obtain complementary information from different modalities, propose a Multi-Modal Multi-Head Attention (MMA) module to compute inter-modality interactions, as illustrated in Fig. 5, which the same as that in the TMA module described in Section IV-B1 following the self-attention paradigm; the MMA will calculate the attention weights indicating the inter-modality correlations, which composes a matrix of weights in
R
m
×
m
at all time steps t ∈ [1, n]).
Claim 14
ZHANG in view of Chen discloses all the elements as stated in Claim 1 and further discloses wherein the inter-modal multi-head attentions discover temporal dependencies between different time steps from the first and second datasets (Chen, Abstract in Page 4171: first introduce the transformer-encoder with a self-attention mechanism and propose a Convolutional Neural Network-Transformer Encoder (CNNTE) framework to model the temporal dependency for single modal affect recognition; further, to effectively consider the complementarity and redundancy between multiple streams, propose a Transformer Encoder with Multi-modal Multi-head Attention (TEMMA) for multi-modal affect recognition; TEMMA allows to progressively and simultaneously refine the inter-modality interactions and intra-modality temporal dependency; the learned multi-modal representations are fed to an Inference Sub-network with fully connected layers to estimate the affective state; Section I in Pages 4171-4172: the transformer, which relies entirely on a self-attention mechanism, demonstrated its effectiveness to draw global dependencies over a sequence; multi-modal transformer frameworks have been built upon the transformer-encoder with a pair-wise cross-modality attention module performing on the low-level descriptors or on the high-level representations; these frameworks considered temporal dynamics and multi-modal dynamic interactions separately; to explore the dynamic interactions between different modalities along with the temporal dependency, an intermediate level multi-modal fusion based on the transformer is proposed; the model constructs multiple feature streams, each consisting of factorized multi-modal features; then the multi-head attention is performed on each feature stream to learn multi-modal dynamic interactions along with temporal dependency; based on the transformer, address the dynamic temporal dependency and multi-modal fusion challenges of continuous affect recognition, as follows: (a) to handle the long-range dependencies, propose to predict affective states by combing one dimensional CNN with the Transformer-Encoder (CNN-TE), wherein the CNN is used to locally aggregate context information, while the transformer-encoder models the long-range dependencies with the attention mechanism; (b) to deal with the dynamic importance of multi-modal feature streams, propose the Multi-modal Multi-head Attention (MMA), which can be easily inserted into the transformer-encoder to promote the dynamic interactions between different modalities and learn their correlations; (c) furthermore, suggest a Transformer-Encoder with Multi-modal Multi-head Attention (TEMMA) and propose a CNN-TEMMA framework, to progressively and simultaneously refine inter-modality interactions and intramodality dependencies, wherein the CNN sub-network extracts and aggregates local temporal features via causal convolution, and the TEMMA sub-network models multi-modal interactions and temporal dynamics, with spatial (audio and visual features) and temporal attention, to learn high-level representations from each modality; and (d) finally, an inference sub-network, with fully connected layers, is used to predict the affective state from the learned multi-modal representations; Section II.B in Page 4173: the authors of [38] extended the transformer of with a pairwise cross-modal attention module to learn representations from different paired multi-modal features; the representations from each pair are further fed into different transformer-encoders to achieve temporal dependency; propose an intermediate level multi-modal fusion scheme built upon the standard Transformer network to efficiently learn intra-modality dynamics and inter-modality dynamics; a novel Multi-modal Multi-head Attention (MMA) module, which learns relationships between several multi-modal sequences, is first presented; then the MMA module is followed by a temporal multi-head attention (TMA) module for each feature stream, composing a CNN and transformer-encoder with multi-modal multi-head attention (CNN-TEMMA) framework; this framework progressively learns multi-modal interactions and temporal dependency; the proposed CNN-TEMMA framework can be used not only for audiovisual continuous affect recognition, but also for other multi-modal feature learning tasks with dynamic sequences as inputs; Section III.B with FIGS. 1-2 in Pages 4174-4175: as shown in Fig. 1, the transformer-encoder is composed of N identical blocks, each consisting of a multi-head attention module followed by a fully connected feed-forward module, which employs multi-head attention (MA) to calculate the temporal dependency (therefore we name it as temporal MA (TMA)); as introduced in [36], the attention function can be described as performing a query on a set of key-values pairs, to generate an output; the output is computed as a weighted sum of values, where the weights is calculated by a compatibility function of the query with the corresponding key; Fig. 2 illustrates the multi-head attention module used in the transformer-encoder, which employs the scaled dot-product attention on the queries (Q), keys (K), and values (V ) under each head; in practice, Q is a set of queries of the whole sequence, packed together in a matrix; the keys and values are also packed together into matrices K and V; Q, K, V are first projected h times in different subspace headi (i ∈ [1, h]) with different learned linear projections
W
i
Q
,
W
i
K
,
W
i
V
of dimensions dk, dk, dv, respectively; then the scaled dot-product attention is performed in parallel on headi of queries, keys, and values; concretely, the scaled dot-product attention computes the dot products (MatMul) of the scaled query and keys; the result is multiplied by a pre-defined attention weights mask, before applying the softmax function to obtain the weights A on the values; such attention block can be mathematically described as eqn. (1), where ∗ represents the Hadamard product, and Mask refers to a bidirectional attention weights mask, which can be regarded as a hard attention mask making the encoder ignore the information from far history and consider the near future; the multi-head attention module (TMA) is given as eqn. (2); apart from the multi-head attention module and the fully connected feed-forward module, the transformer-encoder contains a residual connection followed by a normalization layer (LN); as illustrated in Fig. 1, the transformer-encoder maps the input sequence of low-level descriptors, to an output sequence of high-level representations; Section IV.B in Pages 4175-4176 with FIG. 4 in Page 4176 and FIG. 5 in Page 4171: the Multi-modal Encoder sub-network of Fig. 4, is composed of N stacked encoder blocks, wherein the encoder block comprises a Multi-modal Multi-head Attention (MMA) module followed by the original Multi-head Attention (TMA) of Fig. 1; such architecture not only dynamically enriches each modality with complementary information from the other modalities but also directs TMA on searching Intra-modality dynamics; furthermore, by stacking the encoder blocks, the encoder could progressively refine the inter-modality interactions and intra-modality dynamics to learn high-level semantic representation with temporal contextual information as well as complementary multi-modal information; Temporal Multi-Head Attention (TMA) Module is performed individually on each modality to capture the temporal dependency by following the same principles as described in Section III-B; to obtain complementary information from different modalities, propose a Multi-Modal Multi-Head Attention (MMA) module to compute inter-modality interactions, as illustrated in Fig. 5, which the same as that in the TMA module described in Section IV-B1 following the self-attention paradigm; the MMA will calculate the attention weights indicating the inter-modality correlations, which composes a matrix of weights in
R
m
×
m
at all time steps t ∈ [1, n]).
Claim 15
ZHANG in view of Chen discloses all the elements as stated in Claim 1 and further discloses wherein the first dataset comprises a first input sequence length, and wherein the second dataset comprises a second input sequence length that is different from the first input sequence length (ZHANG, ¶¶ [0014]-[0030] and [0046]-[0047] with FIGS. 1-2: the stock price information refers to the collected stock market closing price data, and the stock price closing sequence data refers to the stock price data after preprocessing the original stock price information. In this invention, the preprocessed stock price closing sequence data is used as the input of the model; social text information refers to information about stocks on various online social platforms, including but not limited to Twitter, Weibo, WeChat, and others; social text information refers to raw text information, while social text data refers to preprocessed text information; in the second step, the text sequence is transformed into a vector feature representation using word vectors, generated according to the following steps: (a) the preprocessed social text data is trained using the word vector model word2vec to learn the word vector representation of each word in the entire text database; the dimension of the word vector is denoted as De; (b) generate vector representations at the social text level; e.g., taking a social media post (Twitter) about a stock as an example, based on the obtained word vectors, average pooling is performed on each dimension of the word vectors of all words in the social media post; the word vector matrix of Nword words in the social text is then subjected to dimensional average pooling to obtain a De dimensional social text representation; (c) generate a vector representation at the day level; e.g., consider social media text related to stocks on a particular day; after obtaining the social text-level vector representations (e.g., Twitter-level representations) of Ntweet social texts (e.g., Twitter) according to the method in step (b) above, for the day level stock text matrix representation with Ntweet*De as the dimension, max, min, and average pooling operations are applied to each dimension of the word vectors to obtain a 3*De stock text representation for that day; realize the transformation of text sequences into vector feature representations using word vectors; in the second step, the operation of performing three-classification processing on the continuous stock price value sequence means that, for the crawled original closing stock price sequence features, if the closing price of the day is higher than the closing price of the previous day, then "+1" is used as the stock price feature of the day; otherwise, "-1" is used as the stock price feature of the day; if the closing price of the day is the same as the closing price of the previous day, then "0" is used as the stock price feature of the day; the continuous stock price sequence features are transformed into a three-class sequence feature taken from {+1, 0, -1}; a recurrent neural network is modeled using external stock prices and target historical price sequence data; input a sequence of M external stocks of length T, [X1, …, XM], where each stock is a vector representation of length T; modeling the social text (Twitter text) dataset using a recurrent neural network refers to: modeling the vectorized social text sequence representation obtained from the preprocessing in the first step using a long short-term memory network; the input is [E1, …, ET], which represents a sequence of T-length Twitter text vectors of the target stock, i.e., the vector representation of social text (e.g., Twitter text) after preprocessing according to the method in the second step, wherein M is different to T) (Chen, Section III.A in Page 4174: the extended Geneva Minimalistic Acoustic Parameter Set (eGeMAPS) features or Mel Filter-bank Cepstrum Coefficient (MFCC) features for audio, and the hand-crafted Local Gabor Binary Patterns from Three Orthogonal Planes (LGBP-TOP) features for face video; use causal convolutions to respect the temporal order during the learning process; in this context, “causal” means that the activations computed for a particular time step do not depend on the activations from the future time steps; Section VI in Page 4181: adopt the hand-crafted audio and visual features (eGeMAPS for audio and LGBP-TOP for video), or extract deep learning features (Vggish for audio and Resnet50 for video), as inputs into the CNN-TE or CNN-TEMMA framework, wherein size of video is different to size of audio).
Claim 16
ZHANG in view of Chen discloses all the elements as stated in Claim 1 and further discloses wherein the feature-level attentions in the first and second modality streams of the encoder produce attention matrices providing explainability of the first and second datasets (Chen, Section II.C in Pages 4173-4174: for enforcing attention on important cues within a sequence, attention mechanisms have been integrated into sequence-to-sequence models; the attention layer can be considered as a weighted-average pooling layer which aggregates contextual information by focusing on the emotional salient information yielding a new representation of the sequence; Section VI in Page 4181: as human observe the changes of temporal affective state from audio and visual cues simultaneously and pay attention to different clues along time, modelling dynamic interactions between audio and visual observations and automatically assigning weights to different cues are of high importance in continuous affect recognition; to model the temporal dynamics and mutual interactions between different modalities for affect recognition, first propose a CNN-Transformer; Encoder (CNN-TE) framework for single modal continuous affect recognition, where CNN is used to aggregate local temporal contextual information, and the transformer-encoder captures long-range dependencies with self-attention; then propose a Multi-modal Multi-head Attention (MMA) module to promote the interactions between different modalities; based on MMA, extend the single modal CNN-TE to a multi-modal CNNTEMMA framework to progressively refine the inter-modality interactions and intra-modality temporal dependency simultaneously; by stacking the TEMMA module, better multi-modal representations can be learned for accurately estimating the affective states; although the CNN-TE and CNN-TEMMA models have been proposed for continuous affect recognition, they could also be used in other multi-modal feature learning applications; adopt the hand-crafted audio and visual features (eGeMAPS for audio and LGBP-TOP for video), or extract deep learning features (Vggish for audio and Resnet50 for video), as inputs into the CNN-TE or CNN-TEMMA framework) (ZHNAG, ¶¶ [0007]-[0011], [0014]-[0026], [0061], [0081]-[0083], [0091]-[0100] and [0123]-[0125]: employ a cross-modal bidirectional attention mechanism to model stock price data and social text, effectively extracting important sequence information; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, allowing them to learn to extract stock price sequences and social text sequences relevant to the prediction target; the word or phrase to be removed is one that is frequently used in English but whose removal does not affect the overall understanding; the key point is to use a long short-term memory network to obtain the hidden state of the memory unit and extract the relationship between stocks and sequences; for the denoised Twitter text, the word vector representation of each word in the text library is learned using the word vector model word2vec; at each time step of the stock price sequence module and the text module are used to extract the state features of the last time step).
Claim 17
ZHANG in view of Chen discloses all the elements as stated in Claim 1 and further discloses in response to the inputting, receiving from the transformer a series of inferred variables related to the target series data, the series of inferred variables being produced in steps with a further predicted value of the series being based off of an earlier predicted value of the series (ZHANG, ¶¶ [0028]-[0065] with FIGS. 1 and 3: recurrent neural networks are used to model the stock price sequence data and the social text dataset, respectively; among them, the modeling of stock price sequence data is as follows: a recurrent neural network is modeled using external stock prices and target historical price sequence data; the core of the model is to use an attention mechanism consisting of an encoder and a decoder: at the encoder end, the attention mechanism is used to select relevant external stocks, and at the decoder end, relevant sequence features are selected for the entire sequence; the state features of the memory unit at each time step input from the encoding end and the input sequence features of the text module are used to select the sequence features related to the predicted value from the entire sequence using an attention mechanism; the weighted sum of the state sequence is used to update the memory cell state at the decoding end together with the historical time series of the target stock; in the third step, modeling the social text (Twitter text) dataset using a recurrent neural network refers to: modeling the vectorized social text sequence representation obtained from the preprocessing in the first step using a long short-term memory network; the input is [E1, …, ET], which represents a sequence of T-length Twitter text vectors of the target stock, i.e., the vector representation of social text (e.g., Twitter text) after preprocessing according to the method in the second step; the weighted sequence sum representation Cd in the stock price sequence module is used to participate in the calculation of text attention weights; this text attention weight can be calculated as a weighted sum of features of the text sequence; Feature Ctext is used to update the state of memory cells/units; in the third step, the use of a bidirectional cross-modal attention mechanism means that, at the decoding end of the stock price network module, the input features of the text [E1, …, ET] are used to help train the sequence attention weights; in the text network module, the weighted sum representation of the hidden states Cd calculated in the stock price module is used to update the text attention weights; therefore, at each moment in the sequence, both modules bidirectionally compute their respective attention weights using cross-modal data from each other; in the fifth step, a network model based on bidirectional cross-modal attention is used to predict the target stock price trend; following the method in step three, obtain the stock price sequence and social text sequence related to the prediction target; i.e., obtain the hidden unit state of the stock price sequence module decoder and the text module at each moment [
h
1
d
e
, …,
h
T
d
e
], [
h
1
t
e
x
t
, …,
h
T
t
e
x
t
]; take the state representation of the last day from the two features in step (1) above and concatenate them to obtain [
h
T
d
e
,
h
T
t
e
x
t
]; prediction is performed using the features of splicing, as follows:
Y
=
ο
(
v
o
W
o
h
T
d
e
;
h
T
t
e
x
t
+
b
o
+
b
v
)
, where sigmoid is used as the activation function σ;
v
o
,
W
o
,
b
o
, and
b
v
are the parameters that need to be trained in the network; ¶¶ [0101]-[0125] with FIG. 3: use the LSTM module in TensorFlow as a basis to perform sequence modeling on the two parts of the sequence data; for stock price sequence data modeling, modeling a recurrent neural network using external stock price sequences and historical price sequences of the target stock; at the encoding end, (a) the input sequence consists of M external stocks of length T, [X1, …, XM], where each stock is a vector representation of length T; (b) calculate attention weights using the input; (c) attention weights are used to select external stock price features relevant to the prediction of stocks; at the decoding end, the state features of the memory unit/cell at each time step input from the input encoder and the input sequence features of the text module are used to select the sequence features in the entire sequence that are relevant to the predicted value using an attention mechanism; attention weights are used to select the relevant memory cell states at the encoder end; for Text sequence data modeling, after obtaining a vectorized text sequence representation through preprocessing, a Long Short-Term Memory (LSTM) network is used to model the text sequence; the initial input is [E1, …, ET], representing a sequence of T-length Twitter text vectors of the target stock, which is the vector representation of the preprocessed Twitter text; the text input features and the weighted sequence representation Cd from the stock price sequence module are used to calculate the text attention weights; during the prediction process, the hidden unit states [
h
1
d
e
, …,
h
T
d
e
], [
h
1
t
e
x
t
, …,
h
T
t
e
x
t
] at each time step of the stock price sequence module and the text module are used to extract the state features of the last time step and concatenate them into [
h
T
d
e
,
h
T
t
e
x
t
] which is then used for prediction:
Y
=
ο
(
v
o
W
o
h
T
d
e
;
h
T
t
e
x
t
+
b
o
+
b
v
)
, where sigmoid is used as the activation function σ;
v
o
,
W
o
,
b
o
, and
b
v
are the parameters that need to be trained in the network).
Claim 18
ZHANG in view of Chen discloses all the elements as stated in Claim 1 and further discloses wherein the inter-modality multi-head attention of the first modality stream uses, as inputs, a keys vector from the first modality stream, a queries vector from the second modality stream, and a values vector from the first modality stream; and wherein the inter-modality multi-head attention of the second modality stream uses, as inputs, a keys vector from the second modality stream, a queries vector from the first modality stream, and a values vector from the second modality stream (Chen, Section IV in Pages 4175-4176 with FIG. 4 in Page 4176 and FIG. 5 in Page 4171: for a given Modalityj, denote by
Q
j
t
the feature vector at time t, and Qj = (
Q
j
1
, …
Q
j
t
, …,
Q
j
n
) the full feature sequence, then Kj = Vj = Qj are the input feature sequences to the TMA module of modality j (Fig. 5); for each modality j, the process can be described as eqn. (3), wherein i ∈ [1,h] denotes the ith head of the total number of h heads, j ∈ [1,m] denotes the jth modality, t ∈ [1,n] denotes the frame, and n the number of frames; the projections are parameter matrices:
W
j
,
i
T
Q
∈
R
d
m
o
d
e
l
×
d
k
,
W
j
,
i
T
K
∈
R
d
m
o
d
e
l
×
d
k
,
W
j
,
i
T
V
∈
R
d
m
o
d
e
l
×
d
v
and
W
j
,
i
T
O
∈
R
h
d
v
×
d
m
o
d
e
l
; the superscripts (T) denotes matrices of the temporal multi-head attention module; propose a MMA module to compute inter-modality interactions, as illustrated in Fig. 5; Let
Q
j
t
the feature vector of modality j at time step t, and Qt = (
Q
1
t
, …
Q
j
t
, …,
Q
m
t
) the multi-modal feature sequence at time step t, then Kt = Vt = Qt are the input multi-modal feature sequences to the MMA module at time step t (Fig. 5), which is the same as that in the TMA module (section IV-B1) following the self-attention paradigm, here the input feature sequences are organized in a multi-modal manner under each time step; the MMA will calculate the attention weights indicating the inter-modality correlations, which composes a matrix of weights in
R
m
×
m
at all time steps t ∈ [1,n]; in practice, for each modality, the same as for TMA, the input is projected into multiple subspaces but with different parameters W(M)Q, W(M)K, W(M)V, wherein the superscript (M) denotes matrices of the MMA module; then, at each time step t, construct the multi-modal queries, keys and values (
Q
i
t
,
K
i
t
,
V
i
t
) under subspace headi followed by the scaled dot-product attention; finally, the output values under each subspace headi are concatenated and linearly projected, resulting in the final values as shown in eqns. (4)-(6)) (ZHANG, ¶¶ [0007]-[0013] with FIGS. 1 and 3: jointly model stock prices and text, and to bidirectionally calculate cross-modal attention weights; employ a cross-modal bidirectional attention mechanism to model stock price data and social text, effectively extracting important sequence information; the stock price prediction method based on a bidirectional cross-modal attention network model; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, allowing them to learn to extract stock price sequences and social text sequences relevant to the prediction target; use a network model based on bidirectional cross-modal attention to predict stock price trends in the target data; ¶ [0039]: the state features of the memory unit at each time step input from the encoding end and the input sequence features of the text module are used to select the sequence features related to the predicted value from the entire sequence using an attention mechanism; ¶¶ [0058], [0060], and [0068] with FIGS.1 and 3: the use of a bidirectional cross-modal attention mechanism means that, at the decoding end of the stock price network module, the input features of the text [E1, …, ET] are used to help train the sequence attention weights; in the text network module, the weighted sum representation of the hidden states Cd calculated in the stock price module is used to update the text attention weights; therefore, at each moment in the sequence, both modules bidirectionally compute their respective attention weights using cross-modal data from each other; a network model based on bidirectional cross-modal attention is used to predict the target stock price trend; the bidirectional cross-modal attention mechanism integrates stock price and text data; its core is the information interaction between the stock price module decoding end and the text module; the method adopted is to utilize the hidden state sequence of the decoding end and the input sequence of the text respectively; ¶¶ [0083] and [0085] with FIG. 3: involve using recurrent neural networks to model the stock price sequence data and the Twitter text dataset, respectively; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, extracting sequence features relevant to the prediction target; predict the stock price trend of the target data based on a network model with bidirectional cross-modal attention).
Claims 4-9 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over ZHANG in view of Chen as applied to Claim 1 above, and further in view of YU (CN 108805087 A, pub. date: 11/13/2018), hereinafter YU.
Claim 4
ZHANG in view of Chen discloses all the elements as stated in Claim 1 and further discloses wherein (ZHANG, ¶¶ [0007]-[0013] with FIG. 1: jointly model stock prices and text, and to bidirectionally calculate cross-modal attention weights; employs a cross-modal bidirectional attention mechanism to model stock price data and social text, effectively extracting important sequence information; the third step involves modeling the stock price sequence data and social text datasets such as Twitter using recurrent neural networks; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, allowing them to learn to extract stock price sequences and social text sequences relevant to the prediction target; ¶¶ [0028]-[0058] and [0068] with FIG. 1 and 3: in the third step, recurrent neural networks are used to model the stock price sequence data and the social text dataset, respectively; among them, the modeling of stock price sequence data is as follows: a recurrent neural network is modeled using external stock prices and target historical price sequence data; the core of the model is to use an attention mechanism consisting of an encoder and a decoder: at the encoder end, the attention mechanism is used to select relevant external stocks, and at the decoder end, relevant sequence features are selected for the entire sequence; use an attention mechanism at the encoding end to select relevant external stocks: (a) input a sequence of M external stocks of length T, [X1, …, XM], where each stock is a vector representation of length T; (b) calculate attention weights using the input; (c) attention weights are used to select external stock price features relevant to the prediction of stocks; this feature is used to update the state value of the memory cell/unit; at the decoding end, relevant sequence features are selected for the entire sequence: (a) the state features of the memory unit at each time step input from the encoding end and the input sequence features of the text module are used to select the sequence features related to the predicted value from the entire sequence using an attention mechanism; wherein, the attention weight is calculated; (b) calculate the weighted sum representation of the state sequence using attention weights; The weighted sum of the state sequence is used to update the memory cell/unit state at the decoding end together with the historical time series of the target stock; in the third step, modeling the social text (Twitter text) dataset using a recurrent neural network refers to: modeling the vectorized social text sequence representation obtained from the preprocessing in the first step using a long short-term memory network; the input is [E1, …, ET], which represents a sequence of T-length Twitter text vectors of the target stock, i.e., the vector representation of social text (e.g., Twitter text) after preprocessing according to the method in the second step; the weighted sequence sum representation Cd in the stock price sequence module is used to participate in the calculation of text attention weights; this text attention weight can be calculated as a weighted sum of features of the text sequence; Feature Ctext is used to update the state of memory cells/units; in the third step, the use of a bidirectional cross-modal attention mechanism means that, at the decoding end of the stock price network module, the input features of the text [E1, …, ET] are used to help train the sequence attention weights; in the text network module, the weighted sum representation of the hidden states Cd calculated in the stock price module is used to update the text attention weights; therefore, at each moment in the sequence, both modules bidirectionally compute their respective attention weights using cross-modal data from each other; in the third step, the bidirectional cross-modal attention mechanism integrates stock price and text data; its core is the information interaction between the stock price module decoding end and the text module; the method adopted is to utilize the hidden state sequence of the decoding end and the input sequence of the text respectively; ¶¶ [0069]-[0072] with FIG. 4: a stock price prediction system utilizing stock price information and social text information, which comprises text and price sequence modeling unit to perform sequence modeling on the stock price data and text data of the input representation, calculate the attention weight of the two data parts using mutual information, and select relevant input representations; ¶ [0083] with FIG. 1: the third step involves using recurrent neural networks to model the stock price sequence data and the Twitter text dataset, respectively; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, extracting sequence features relevant to the prediction target; ¶¶ [0086]-[0089] with FIG. 4: a stock price prediction system comprising: text and price series modeling unit, where (a) a long short-term memory network is used to perform sequence modeling on the input representations of stock price data and text data; and (b) a bidirectional attention mechanism is used to select relevant input representations by calculating the attention weights of the two data parts using mutual information; ¶¶ [0101]-[0121] with FIGS. 1 and 3: use the LSTM module in TensorFlow as a basis to perform sequence modeling on the two parts of the sequence data; for stock price sequence data modeling, modeling a recurrent neural network using external stock price sequences and historical price sequences of the target stock; at the encoding end, (a) the input sequence consists of M external stocks of length T, [X1, …, XM], where each stock is a vector representation of length T; (b) calculate attention weights using the input; (c) attention weights are used to select external stock price features relevant to the prediction of stocks; at the decoding end, the state features of the memory unit/cell at each time step input from the input encoder and the input sequence features of the text module are used to select the sequence features in the entire sequence that are relevant to the predicted value using an attention mechanism; attention weights are used to select the relevant memory cell states at the encoder end; for Text sequence data modeling, after obtaining a vectorized text sequence representation through preprocessing, a Long Short-Term Memory (LSTM) network is used to model the text sequence; the initial input is [E1, …, ET], representing a sequence of T-length Twitter text vectors of the target stock, which is the vector representation of the preprocessed Twitter text; the text input features and the weighted sequence representation Cd from the stock price sequence module are used to calculate the text attention weights).
ZHANG in view of Chen fails to explicitly disclose wherein the target series data are input into the decoder, the decoder performs multi-head attention on the target series data, and a signal from the multi-head attention of the decoder is combined with the output from the encoder.
YU teaches a system and a method using attention mechanism (YU, ¶ [0039]), wherein the target series data are input into the decoder, the decoder performs multi-head attention on the target series data, and a signal from the multi-head attention of the decoder is combined with the output from the encoder (YU, ¶ [0039]: the multi-turn dialogue semantic understanding subsystem adds an emotion recognition attention mechanism to the input utterance of the current turn on the basis of the traditional seq2seq language generation model, and incorporates emotion tracking from previous turns of dialogue in the time sequence into the dialogue management; each utterance spoken by the current user is input into a bidirectional LSTM encoder, and then the input of the currently identified different emotional states is merged with the encoder output of the previously generated user utterance, and input together into the decoder; in this way, the decoder has both the user's utterance and the current emotion, and the system dialogue response generated afterward is a personalized output specific to the current user's emotional state; the emotion-aware information state update (ISU) strategy updates the dialogue state at any moment when new information is available; each update of the dialogue state is deterministic, and for the same system state, the same system behavior, and the same current user emotional state at the previous moment, the same current system state will inevitably be generated; ¶¶ [0148]-[0155] and [0166] with FIGS. 17-18: propose to focus on improving dialogue management, strengthening language understanding and attention mechanisms for emotional words, which can effectively grasp the basic semantics and capture emotions in multi-turn dialogues; an attention mechanism for emotion recognition is added to the input utterance of the current round on the basis of the traditional seq2seq language generation model; emotion tracking in the previous rounds of dialogue is added to the dialogue management; the input utterance of the current round is based on the traditional seq2seq language generation model, with the addition of an attention mechanism for emotion recognition; each utterance spoken by the current user is input into a bidirectional LSTM encoder; unlike traditional language generation models, attention is added to the emotion in the current sentence; next, the current input of different emotional states is merged with the encoder output of the user's speech generated earlier and input into the decoder; in this way, the decoder has both the user's speech and the current emotion, and the system dialogue response generated afterward is a personalized output specific to the current user's emotional state; emotion recognition in multi-turn dialogue is a simple method for updating the dialogue state: the Sentiment Aware Information State Update (ISU) strategy; the SAISU strategy updates the conversation state whenever new information is available; specifically, the conversation state is updated whenever new information is generated by the user, the system, or any participant in the conversation; this update is based on previous rounds of emotion perception; Figure 18 shows that the dialogue state st+1 at time t+1 depends on the state st at the previous time t, the system behavior at the previous time t, and the user behavior and emotion ot+1 at the current time t+1; for the same system state and behavior in the previous moment, and the same user emotional state in the current moment, the same system state in the current moment will necessarily be generated; ¶¶ [0166]-[0168] with FIG. 19: a deep neural network is used to encode, deeply correlate and understand information from multiple single modalities and then make a comprehensive judgment; the overall architecture considers that emotion recognition is carried out on a continuous timeline, and makes a judgment on the current point in time based on all related facial expressions, actions, text, speech and physiological data; therefore, this method was based on the classic seq2seq neural network; the main idea behind Seq2Seq is to map an input sequence to an output sequence using a deep neural network model (commonly LSTM, Long Short Memory, a type of recurrent neural network); this process consists of two stages: encoding the input and decoding the output; when the basic seq2seq model is applied to emotion recognition analysis based on a continuous time axis, it needs unique and innovative variations to better solve specific problems; i emotion recognition, in addition to the problems that the usual seq2seq model needs to handle, the following key characteristics also need to be considered: (a) the relationship between different time points of multiple unimodalities; (b) the inherent influence and relationship between multimodalities at the same time point; and (c) the overall recognition and identification of emotions by integrating multimodalities; these problems have not been solved in existing technologies; specifically, the model first includes 5 recurrent neural networks (RNN); in practical systems, use long-short term memory (LSTM), a representative of RNNs, wherein each RNN organizes the intermediate neural network representations of each unimodal emotion understanding in a time sequence; at each time point (a blue bar in Figure 19), a neural network unit comes from the output of the corresponding time point of the intermediate layer of the neural network of the single-modal subsystem described earlier; the output of each RNN at a single time point (a blue bar in Figure 19) is fed into the multimodal fusion association judgment RNN; therefore, each time point of a multimodal RNN aggregates the neural network outputs of each unimodal RNN at the current time point; after integrating the multimodal data, the output at each time point is the final sentiment judgment result at that time point (orange arrow in Figure 19))
ZHAN and YU are analogous art because they are from the same field of endeavor, a system and method using attention mechanism. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of YU to ZHANG in view of Chen. Motivation for doing so would provide u.
Claim 5
ZHANG in view of Chen and YU discloses all the elements as stated in Claim 4 and further discloses wherein the output from the encoder comprises a first output signal from the first modality stream and a second output signal from the second modality stream, and wherein the signal from the multi-head attention of the decoder is combined separately with the first output signal and the second output signal for separate analysis of dependencies between: the target series data and the first dataset and the target series data and the second dataset (Chen, Section IV in Pages 4175-4176 with FIG. 4 in Page 4176 and FIG. 5 in Page 4171: for a given Modalityj, denote by
Q
j
t
the feature vector at time t, and Qj = (
Q
j
1
, …
Q
j
t
, …,
Q
j
n
) the full feature sequence, then Kj = Vj = Qj are the input feature sequences to the TMA module of modality j (Fig. 5); for each modality j, the process can be described as eqn. (3), wherein i ∈ [1,h] denotes the ith head of the total number of h heads, j ∈ [1,m] denotes the jth modality, t ∈ [1,n] denotes the frame, and n the number of frames; the projections are parameter matrices:
W
j
,
i
T
Q
∈
R
d
m
o
d
e
l
×
d
k
,
W
j
,
i
T
K
∈
R
d
m
o
d
e
l
×
d
k
,
W
j
,
i
T
V
∈
R
d
m
o
d
e
l
×
d
v
and
W
j
,
i
T
O
∈
R
h
d
v
×
d
m
o
d
e
l
; the superscripts (T) denotes matrices of the temporal multi-head attention module; propose a MMA module to compute inter-modality interactions, as illustrated in Fig. 5; Let
Q
j
t
the feature vector of modality j at time step t, and Qt = (
Q
1
t
, …
Q
j
t
, …,
Q
m
t
) the multi-modal feature sequence at time step t, then Kt = Vt = Qt are the input multi-modal feature sequences to the MMA module at time step t (Fig. 5), which is the same as that in the TMA module (section IV-B1) following the self-attention paradigm, here the input feature sequences are organized in a multi-modal manner under each time step; the MMA will calculate the attention weights indicating the inter-modality correlations, which composes a matrix of weights in
R
m
×
m
at all time steps t ∈ [1,n]; in practice, for each modality, the same as for TMA, the input is projected into multiple subspaces but with different parameters W(M)Q, W(M)K, W(M)V, wherein the superscript (M) denotes matrices of the MMA module; then, at each time step t, construct the multi-modal queries, keys and values (
Q
i
t
,
K
i
t
,
V
i
t
) under subspace headi followed by the scaled dot-product attention; finally, the output values under each subspace headi are concatenated and linearly projected, resulting in the final values as shown in eqns. (4)-(6)) (ZHANG, ¶¶ [0007]-[0013] with FIGS. 1 and 3: jointly model stock prices and text, and to bidirectionally calculate cross-modal attention weights; employ a cross-modal bidirectional attention mechanism to model stock price data and social text, effectively extracting important sequence information; the stock price prediction method based on a bidirectional cross-modal attention network model; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, allowing them to learn to extract stock price sequences and social text sequences relevant to the prediction target; use a network model based on bidirectional cross-modal attention to predict stock price trends in the target data; ¶ [0039]: the state features of the memory unit at each time step input from the encoding end and the input sequence features of the text module are used to select the sequence features related to the predicted value from the entire sequence using an attention mechanism; ¶¶ [0058], [0060], and [0068] with FIGS.1 and 3: the use of a bidirectional cross-modal attention mechanism means that, at the decoding end of the stock price network module, the input features of the text [E1, …, ET] are used to help train the sequence attention weights; in the text network module, the weighted sum representation of the hidden states Cd calculated in the stock price module is used to update the text attention weights; therefore, at each moment in the sequence, both modules bidirectionally compute their respective attention weights using cross-modal data from each other; a network model based on bidirectional cross-modal attention is used to predict the target stock price trend; the bidirectional cross-modal attention mechanism integrates stock price and text data; its core is the information interaction between the stock price module decoding end and the text module; the method adopted is to utilize the hidden state sequence of the decoding end and the input sequence of the text respectively; ¶¶ [0083] and [0085] with FIG. 3: involve using recurrent neural networks to model the stock price sequence data and the Twitter text dataset, respectively; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, extracting sequence features relevant to the prediction target; predict the stock price trend of the target data based on a network model with bidirectional cross-modal attention).
Claim 6
ZHANG in view of Chen and YU discloses all the elements as stated in Claim 5 and further discloses wherein the separate analysis of dependencies occurs in a first target cross-attention mechanism and in a second target cross-attention mechanism, and wherein respective outputs from the first and second target cross-attention mechanisms of the decoder are combined to produce the inferred variable related to the target series data (ZHANG, ¶¶ [0028]-[0065] with FIGS. 1 and 3: recurrent neural networks are used to model the stock price sequence data and the social text dataset, respectively; among them, the modeling of stock price sequence data is as follows: a recurrent neural network is modeled using external stock prices and target historical price sequence data; the core of the model is to use an attention mechanism consisting of an encoder and a decoder: at the encoder end, the attention mechanism is used to select relevant external stocks, and at the decoder end, relevant sequence features are selected for the entire sequence; the state features of the memory unit at each time step input from the encoding end and the input sequence features of the text module are used to select the sequence features related to the predicted value from the entire sequence using an attention mechanism; the weighted sum of the state sequence is used to update the memory cell state at the decoding end together with the historical time series of the target stock; in the third step, modeling the social text (Twitter text) dataset using a recurrent neural network refers to: modeling the vectorized social text sequence representation obtained from the preprocessing in the first step using a long short-term memory network; the input is [E1, …, ET], which represents a sequence of T-length Twitter text vectors of the target stock, i.e., the vector representation of social text (e.g., Twitter text) after preprocessing according to the method in the second step; the weighted sequence sum representation Cd in the stock price sequence module is used to participate in the calculation of text attention weights; this text attention weight can be calculated as a weighted sum of features of the text sequence; Feature Ctext is used to update the state of memory cells/units; in the third step, the use of a bidirectional cross-modal attention mechanism means that, at the decoding end of the stock price network module, the input features of the text [E1, …, ET] are used to help train the sequence attention weights; in the text network module, the weighted sum representation of the hidden states Cd calculated in the stock price module is used to update the text attention weights; therefore, at each moment in the sequence, both modules bidirectionally compute their respective attention weights using cross-modal data from each other; in the fifth step, a network model based on bidirectional cross-modal attention is used to predict the target stock price trend; following the method in step three, obtain the stock price sequence and social text sequence related to the prediction target; i.e., obtain the hidden unit state of the stock price sequence module decoder and the text module at each moment [
h
1
d
e
, …,
h
T
d
e
], [
h
1
t
e
x
t
, …,
h
T
t
e
x
t
]; take the state representation of the last day from the two features in step (1) above and concatenate them to obtain [
h
T
d
e
,
h
T
t
e
x
t
]; prediction is performed using the features of splicing, as follows:
Y
=
ο
(
v
o
W
o
h
T
d
e
;
h
T
t
e
x
t
+
b
o
+
b
v
)
, where sigmoid is used as the activation function σ;
v
o
,
W
o
,
b
o
, and
b
v
are the parameters that need to be trained in the network; ¶¶ [0101]-[0125] with FIG. 3: use the LSTM module in TensorFlow as a basis to perform sequence modeling on the two parts of the sequence data; for stock price sequence data modeling, modeling a recurrent neural network using external stock price sequences and historical price sequences of the target stock; at the encoding end, (a) the input sequence consists of M external stocks of length T, [X1, …, XM], where each stock is a vector representation of length T; (b) calculate attention weights using the input; (c) attention weights are used to select external stock price features relevant to the prediction of stocks; at the decoding end, the state features of the memory unit/cell at each time step input from the input encoder and the input sequence features of the text module are used to select the sequence features in the entire sequence that are relevant to the predicted value using an attention mechanism; attention weights are used to select the relevant memory cell states at the encoder end; for Text sequence data modeling, after obtaining a vectorized text sequence representation through preprocessing, a Long Short-Term Memory (LSTM) network is used to model the text sequence; the initial input is [E1, …, ET], representing a sequence of T-length Twitter text vectors of the target stock, which is the vector representation of the preprocessed Twitter text; the text input features and the weighted sequence representation Cd from the stock price sequence module are used to calculate the text attention weights; during the prediction process, the hidden unit states [
h
1
d
e
, …,
h
T
d
e
], [
h
1
t
e
x
t
, …,
h
T
t
e
x
t
] at each time step of the stock price sequence module and the text module are used to extract the state features of the last time step and concatenate them into [
h
T
d
e
,
h
T
t
e
x
t
] which is then used for prediction:
Y
=
ο
(
v
o
W
o
h
T
d
e
;
h
T
t
e
x
t
+
b
o
+
b
v
)
, where sigmoid is used as the activation function σ;
v
o
,
W
o
,
b
o
, and
b
v
are the parameters that need to be trained in the network).
Claim 7
ZHANG in view of Chen and YU discloses all the elements as stated in Claim 4 and further discloses wherein the output from the encoder is used as key vector and a value vector for a cross-attention layer in the decoder and a query vector for the cross-attention layer comes from the target series data via the decoder (Chen, Section IV in Pages 4175-4176 with FIG. 4 in Page 4176 and FIG. 5 in Page 4171: for a given Modalityj, denote by
Q
j
t
the feature vector at time t, and Qj = (
Q
j
1
, …
Q
j
t
, …,
Q
j
n
) the full feature sequence, then Kj = Vj = Qj are the input feature sequences to the TMA module of modality j (Fig. 5); for each modality j, the process can be described as eqn. (3), wherein i ∈ [1,h] denotes the ith head of the total number of h heads, j ∈ [1,m] denotes the jth modality, t ∈ [1,n] denotes the frame, and n the number of frames; the projections are parameter matrices:
W
j
,
i
T
Q
∈
R
d
m
o
d
e
l
×
d
k
,
W
j
,
i
T
K
∈
R
d
m
o
d
e
l
×
d
k
,
W
j
,
i
T
V
∈
R
d
m
o
d
e
l
×
d
v
and
W
j
,
i
T
O
∈
R
h
d
v
×
d
m
o
d
e
l
; the superscripts (T) denotes matrices of the temporal multi-head attention module; propose a MMA module to compute inter-modality interactions, as illustrated in Fig. 5; Let
Q
j
t
the feature vector of modality j at time step t, and Qt = (
Q
1
t
, …
Q
j
t
, …,
Q
m
t
) the multi-modal feature sequence at time step t, then Kt = Vt = Qt are the input multi-modal feature sequences to the MMA module at time step t (Fig. 5), which is the same as that in the TMA module (section IV-B1) following the self-attention paradigm, here the input feature sequences are organized in a multi-modal manner under each time step; the MMA will calculate the attention weights indicating the inter-modality correlations, which composes a matrix of weights in
R
m
×
m
at all time steps t ∈ [1,n]; in practice, for each modality, the same as for TMA, the input is projected into multiple subspaces but with different parameters W(M)Q, W(M)K, W(M)V, wherein the superscript (M) denotes matrices of the MMA module; then, at each time step t, construct the multi-modal queries, keys and values (
Q
i
t
,
K
i
t
,
V
i
t
) under subspace headi followed by the scaled dot-product attention; finally, the output values under each subspace headi are concatenated and linearly projected, resulting in the final values as shown in eqns. (4)-(6)).
Claim 8
ZHANG in view of Chen and YU discloses all the elements as stated in Claim 4 and further discloses wherein the decoder ascertains keys, values, and queries from the target series data and inputs the keys, the values, and the queries into the multi-head attention (Chen, Section IV in Pages 4175-4176 with FIG. 4 in Page 4176 and FIG. 5 in Page 4171: for a given Modalityj, denote by
Q
j
t
the feature vector at time t, and Qj = (
Q
j
1
, …
Q
j
t
, …,
Q
j
n
) the full feature sequence, then Kj = Vj = Qj are the input feature sequences to the TMA module of modality j (Fig. 5); for each modality j, the process can be described as eqn. (3), wherein i ∈ [1,h] denotes the ith head of the total number of h heads, j ∈ [1,m] denotes the jth modality, t ∈ [1,n] denotes the frame, and n the number of frames; the projections are parameter matrices:
W
j
,
i
T
Q
∈
R
d
m
o
d
e
l
×
d
k
,
W
j
,
i
T
K
∈
R
d
m
o
d
e
l
×
d
k
,
W
j
,
i
T
V
∈
R
d
m
o
d
e
l
×
d
v
and
W
j
,
i
T
O
∈
R
h
d
v
×
d
m
o
d
e
l
; the superscripts (T) denotes matrices of the temporal multi-head attention module; propose a MMA module to compute inter-modality interactions, as illustrated in Fig. 5; Let
Q
j
t
the feature vector of modality j at time step t, and Qt = (
Q
1
t
, …
Q
j
t
, …,
Q
m
t
) the multi-modal feature sequence at time step t, then Kt = Vt = Qt are the input multi-modal feature sequences to the MMA module at time step t (Fig. 5), which is the same as that in the TMA module (section IV-B1) following the self-attention paradigm, here the input feature sequences are organized in a multi-modal manner under each time step; the MMA will calculate the attention weights indicating the inter-modality correlations, which composes a matrix of weights in
R
m
×
m
at all time steps t ∈ [1,n]; in practice, for each modality, the same as for TMA, the input is projected into multiple subspaces but with different parameters W(M)Q, W(M)K, W(M)V, wherein the superscript (M) denotes matrices of the MMA module; then, at each time step t, construct the multi-modal queries, keys and values (
Q
i
t
,
K
i
t
,
V
i
t
) under subspace headi followed by the scaled dot-product attention; finally, the output values under each subspace headi are concatenated and linearly projected, resulting in the final values as shown in eqns. (4)-(6)).
Claim 9
ZHANG in view of Chen and YU n discloses all the elements as stated in Claim 4 and further discloses wherein the multi-head attention of the decoder is masked-multi-head attention (Chen, Section III with FIGS. 1-3 in Pages 4174-4175: introduce our proposed uni-modal affect recognition framework, which combines a CNN and a Transformer-Encoder (CNN-TE), as illustrated in Fig. 1; the extracted feature vectors are fed into the 1-dimensional temporal convolutional network for aggregating local temporal context information; the output of the 1-D CNN is input into the transformer-encoder to achieve long-range dependencies with dynamic attention weights; finally, an inference sub-network, with fully connected layers, is used to estimate the valence or arousal dimension of the affective state; a 1-D temporal convolution network is adopted to encode the temporal information from the input feature sequence; in the model, use causal convolutions to respect the temporal order during the learning process, wherein "causal" means that the activations computed for a particular time step do not depend on the activations from the future time steps; since the transformer-encoder contains neither recurrence nor convolution, to make use of the time step order of the feature vectors within the sequence, add the position information to the output of the 1-D temporal convolutional network; as shown in Fig. 1, the transformer-encoder is composed of N identical blocks, each consisting of a multi-head attention module followed by a fully connected feed-forward module, which employs multi-head attention (MA) to calculate the temporal dependency (therefore we name it as temporal MA (TMA)); as introduced in [36], the attention function can be described as performing a query on a set of key-values pairs, to generate an output; the output is computed as a weighted sum of values, where the weights is calculated by a compatibility function of the query with the corresponding key; Fig. 2 illustrates the multi-head attention module used in the transformer-encoder, which employs the scaled dot-product attention on the queries (Q), keys (K), and values (V ) under each head; in practice, Q is a set of queries of the whole sequence, packed together in a matrix; the keys and values are also packed together into matrices K and V; Q, K, V are first projected h times in different subspace headi (i ∈ [1, h]) with different learned linear projections
W
i
Q
,
W
i
K
,
W
i
V
of dimensions dk, dk, dv, respectively; then the scaled dot-product attention is performed in parallel on headi of queries, keys, and values; concretely, the scaled dot-product attention computes the dot products (MatMul) of the scaled query and keys; the result is multiplied by a pre-defined attention weights mask, before applying the softmax function to obtain the weights A on the values; such attention block can be mathematically described as eqn. (1), where ∗ represents the Hadamard product, and Mask refers to a bidirectional attention weights mask, which can be regarded as a hard attention mask making the encoder ignore the information from far history and consider the near future; the multi-head attention module (TMA) is given as eqn. (2); apart from the multi-head attention module and the fully connected feed-forward module, the transformer-encoder contains a residual connection followed by a normalization layer (LN); as illustrated in Fig. 1, the transformer-encoder maps the input sequence of low-level descriptors, to an output sequence of high-level representations; the learned high-level representations are fed to an Inference sub-network to estimate the arousal or valence dimension of the affective state, which is composed of two fully connected layers with a non-linear activation layer in between; the number of nodes of the last fully connected layer is half of the dimension of the high-level representations).
Claim 11
ZHANG in view of Chen and YU discloses all the elements as stated in Claim 9 and further discloses wherein the feature-level attentions in the first and second modality streams of the encoder generate first weights for the first modality stream and second weights, different from the first weights, for the second modality stream, respectively, based on same time steps from the first and second datasets (ZHANG, ¶¶ [0007]-[0013] with FIG. 1: jointly model stock prices and text, and to bidirectionally calculate cross-modal attention weights; employs a cross-modal bidirectional attention mechanism to model stock price data and social text, effectively extracting important sequence information; the third step involves modeling the stock price sequence data and social text datasets such as Twitter using recurrent neural networks; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, allowing them to learn to extract stock price sequences and social text sequences relevant to the prediction target; ¶¶ [0028]-[0058] and [0068] with FIG. 1 and 3: in the third step, recurrent neural networks are used to model the stock price sequence data and the social text dataset, respectively; among them, the modeling of stock price sequence data is as follows: a recurrent neural network is modeled using external stock prices and target historical price sequence data; the core of the model is to use an attention mechanism consisting of an encoder and a decoder: at the encoder end, the attention mechanism is used to select relevant external stocks, and at the decoder end, relevant sequence features are selected for the entire sequence; use an attention mechanism at the encoding end to select relevant external stocks: (a) input a sequence of M external stocks of length T, [X1, …, XM], where each stock is a vector representation of length T; (b) calculate attention weights using the input; (c) attention weights are used to select external stock price features relevant to the prediction of stocks; this feature is used to update the state value of the memory cell/unit; at the decoding end, relevant sequence features are selected for the entire sequence: (a) the state features of the memory unit at each time step input from the encoding end and the input sequence features of the text module are used to select the sequence features related to the predicted value from the entire sequence using an attention mechanism; wherein, the attention weight is calculated; (b) calculate the weighted sum representation of the state sequence using attention weights; The weighted sum of the state sequence is used to update the memory cell/unit state at the decoding end together with the historical time series of the target stock; in the third step, modeling the social text (Twitter text) dataset using a recurrent neural network refers to: modeling the vectorized social text sequence representation obtained from the preprocessing in the first step using a long short-term memory network; the input is [E1, …, ET], which represents a sequence of T-length Twitter text vectors of the target stock, i.e., the vector representation of social text (e.g., Twitter text) after preprocessing according to the method in the second step; the weighted sequence sum representation Cd in the stock price sequence module is used to participate in the calculation of text attention weights; this text attention weight can be calculated as a weighted sum of features of the text sequence; Feature Ctext is used to update the state of memory cells/units; in the third step, the use of a bidirectional cross-modal attention mechanism means that, at the decoding end of the stock price network module, the input features of the text [E1, …, ET] are used to help train the sequence attention weights; in the text network module, the weighted sum representation of the hidden states Cd calculated in the stock price module is used to update the text attention weights; therefore, at each moment in the sequence, both modules bidirectionally compute their respective attention weights using cross-modal data from each other; in the third step, the bidirectional cross-modal attention mechanism integrates stock price and text data; its core is the information interaction between the stock price module decoding end and the text module; the method adopted is to utilize the hidden state sequence of the decoding end and the input sequence of the text respectively; ¶¶ [0069]-[0072] with FIG. 4: a stock price prediction system utilizing stock price information and social text information, which comprises text and price sequence modeling unit to perform sequence modeling on the stock price data and text data of the input representation, calculate the attention weight of the two data parts using mutual information, and select relevant input representations; ¶ [0083] with FIG. 1: the third step involves using recurrent neural networks to model the stock price sequence data and the Twitter text dataset, respectively; a bidirectional cross-modal attention mechanism is then used to fuse the two modules, extracting sequence features relevant to the prediction target; ¶¶ [0086]-[0089] with FIG. 4: a stock price prediction system comprising: text and price series modeling unit, where (a) a long short-term memory network is used to perform sequence modeling on the input representations of stock price data and text data; and (b) a bidirectional attention mechanism is used to select relevant input representations by calculating the attention weights of the two data parts using mutual information; ¶¶ [0101]-[0121] with FIGS. 1 and 3: use the LSTM module in TensorFlow as a basis to perform sequence modeling on the two parts of the sequence data; for stock price sequence data modeling, modeling a recurrent neural network using external stock price sequences and historical price sequences of the target stock; at the encoding end, (a) the input sequence consists of M external stocks of length T, [X1, …, XM], where each stock is a vector representation of length T; (b) calculate attention weights using the input; (c) attention weights are used to select external stock price features relevant to the prediction of stocks; at the decoding end, the state features of the memory unit/cell at each time step input from the input encoder and the input sequence features of the text module are used to select the sequence features in the entire sequence that are relevant to the predicted value using an attention mechanism; attention weights are used to select the relevant memory cell states at the encoder end; for Text sequence data modeling, after obtaining a vectorized text sequence representation through preprocessing, a Long Short-Term Memory (LSTM) network is used to model the text sequence; the initial input is [E1, …, ET], representing a sequence of T-length Twitter text vectors of the target stock, which is the vector representation of the preprocessed Twitter text; the text input features and the weighted sequence representation Cd from the stock price sequence module are used to calculate the text attention weights).
Response to Arguments
Applicant's arguments filed on 06/19/2026 have been fully considered but they are not persuasive.
Applicant argues on Pages 9-10 of the Remarks with respect to Claims 1 and 19-20 that Chen does not teach the claim 1 limitation that "each of the first and second modality streams respectively perform[s] feature-level attention, intra-modal multi-head attention, and inter-modal multi-head attention." because a 1D convolutional embedding is not feature-level attention, and Chen does not disclose that each modality stream performs feature-level attention.
In response, examiner respectfully disagrees. According to the specification of the claimed invention described in ¶ [0044] with FIG. 7 that the "feature attention layer" is also use one dimensional convolutional layer to process input data, and therefore, Chen's "Input Embedding Sub-network" in FIG.4, using 1-D CNN for extracting "local temporal features" from input data, is INDEED equivalent to "feature attention layer" of the claimed invention. According to FIG. 4 (shown below) with Section IV of Chen, Video stream (blue line) and Audio stream (red line) respectively performing feature-level attention (Input Embedding Sub-network for extracting local temporal features), intra-modal multi-head attention (TMA, i.e., data is not cross-over), and inter-modal multi-head attention (MMA, i.e., data is cross-over). Therefore, Chen DOES INDEED teach "each of the first and second modality streams respectively perform[s] feature-level attention, intra-modal multi-head attention, and inter-modal multi-head attention" as recited in the claim.
PNG
media_image1.png
336
808
media_image1.png
Greyscale
Applicant further argues on Page 10 of the Remarks regarding Claim 2 that Chen does not appear to teach or suggest temporal features identified by intra-modal multi-head attentions being input into inter-modal multi-head attention.
In response, examiner respectfully disagrees. Although Chen teaches the sequence in reverse order (i.e., output of inter-modal multi-head attention (MMA, i.e., data is cross-over) being input to intra-modal multi-head attention (TMA, i.e., data is not cross-over), however, primary reference ZHANG teaches in FIG. 3 (shown below) that stock price sequence and text sequence are parallel processed separately (i.e., data not cross-over) first and then their output being input to next module with data cross-over; i.e., output of intra-modal is input into inter-modal as claimed. Since the operation order of "inter-modal multi-head attention" and "intra-modal multi-head attention" is simply a design choice (see also FIG. 1 with Section 3 of refernce Curto et al. described in Conclusion section, wherein output of self-attention encoder being input of cross-attention encoder and output of cross-attention encoder being input to another self-attention encode), it is obvious to ordinary skill person in the art to choose a proper order or design for a particular problem. Therefore, the combination of ZHANG and Chen teaches all the limitations as recited in Claim 2.
PNG
media_image2.png
723
1391
media_image2.png
Greyscale
Applicant further argues on Pages 10-11 of the Remarks with respect to Claim 18 that "the inter-modality multi-head attention of the first modality stream uses, as inputs, a keys vector from the first modality stream, a queries vector from the second modality stream, and a values vector from the first modality stream; and wherein the inter-modality multi-head attention of the second modality stream uses, as inputs, a keys vector from the second modality stream, a queries vector from the first modality stream, and a values vector from the second modality stream"; i.e., for a given modality stream, the keys and values come from that stream while the queries come from the other stream.
In response, examiner respectfully disagrees. Chen in FIG. 4 also teaches Video (blue) and Audio (red) streams are crossed-over in MMA, and according to Equations (4)-(6) with FIG.5 that Q, K, V are all crossed-over and mixed together with different weights in MMA. In other word, the claimed invention in Claim 18 with FIG. 3, which describes only Q are crossed-over between text stream (Qtxt) and time stream (Qts) and both K and V are not crossed-over, is just one of scenarios taught by Chen with special weights assigned (e.g., for m=2, WQ(i, j)=0 when i=j; otherwise, WQ(i, j)=1 (i.e., Q is crossed-over); WK(i, j)=1 when i=j; otherwise, WK(i, j)=0 (i.e., K is not crossed-over); WV(i, j)=1 when i=j; otherwise, WV(i, j)=0 (i.e., V is not crossed-over). Similarly, FIG. 4 in reference Lu and FIG.1 in Curto (see Conclusion section) is another scenario with K and V are crossed-over by Q is not crossed-over. Therefore, Chen DOES teach all possible scenarios including one of scenarios cited in Claim 18.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Lu et al. ("MTCA: A Multimodal Summarization Model Based on Two-Stream Cross Attention", 2022 2nd International Conference on Computer Science, Electronic Information Engineering and Intelligent Control Technology (CEI), Sep. 23-25, 2022, pp. 594-601) discloses in Abstract and Section I of Page 594 that (1) propose MTCA which composed of a pre-trained feature extractor, a text encoder, an image encoder, a two-stream cross attention fusion module and a summary decoder to complete the multimodal summarization task; (2) this framework integrates three sub tasks: extractive summarization, abstractive summarization and image selection; (3) at the same time, we introduce a residual network to alleviate the problem of modality-bias; (4) multimodal summarization integrating text and image improves users' satisfaction by providing users with high-quality abstractive summarization with texts and images; (5) the visual modality information in multimodal data can even be used to improve the quality of text summarization, which is helpful for text understanding; (6) the image information interspersed in the text can effectively help users understand the content, enrich the sense, and optimize the experience; (7) therefore, how to reasonably combine the text and image information to generate a multimodal output summary and present it to users is a valuable research direction; (8) at the same time, this technology can be widely used in news push, cross-border e-commerce, automatic generation of product description and other fields, and has rich application prospects in real life. Lu further discloses in Section III with FIGS. 2-4 of Pages 595-597 that (1) as shown in Figure 2, this paper proposes a multimodal cross attention model (MTCA) composed of a pre-trained text and image feature extractor, a two-stream cross attention fusion module and a summary decoder to complete the multimodal summarization task that integrates three subtasks: extractive summarization, abstractive summarization and image selection; (2) because text modal data contains more semantic information, text modal data is more important in MSMO tasks, therefore, a residual connection is applied in the model to enhance the text modal information and alleviate the modality-bias problem; (3) transformer has been applied to the multimodal field; multimodal fusion based on transformer can be divided into one stream model and two stream model; (4) the one stream model uses a transformer to fully interact with multimodal information at the beginning; (5) the two stream model uses independent transformer coding for different modals, and then realizes the fusion between different modals through the co-attention mechanism; (6) the two stream model can adapt to the independent processing needs of different modals; (7) ViLBERT (Vision-and-Language BERT)[15] proved that the performance of the two stream model is better than that of the one stream model; (8) the two stream structure fully learns the features of each modal encoding the visual part and the text part respectively, and then cross encoding; (9) compared with the one stream structure, the two stream structure has better ability to efficiently extract feature during the learning procedure, which is similar to the obtain the manifold feature based on one extracted feature of different modalities; (10) this paper adopts the two stream attention mechanism, and uses the two stream attention network in ViLBERT for reference; (11) in the encoding stage, the image and text are first extracted by their corresponding encoders, and then interact with each other through the multimodal cross attention fusion module; finally, the text summary and the selected image are output; (12) as shown in Figure 3, in this paper, the text encoder uses a pre-trained BART[16] model for text encoding, and outputs the text vector after multi-layer transformer stacked neural network encoding; (13) the image extracts visual features through ResNet-152[17]; (14) in this paper, we use BART as our text encoder; (15) we select the best performing ResNet-152 as the image encoder to extract the 2048 dimensional global features of the average pooling layer at the last layer of the model, and obtain the image characterization matrix through the full connection layer; (16) when the two modals are encoded separately, their output passes through a co-attention module, where the information between different modals is fused, which is expressed as text attention under image conditions in the visual stream and image attention under text conditions in the text stream; (17) the internal composition of the module is shown in Figure 4; (18) this module is based on the structure of transformer, but in the self-attention mechanism, each modal uses its own query to calculate attention with the value and key of another modal, so as to fuse the information between different modals; (19) the inputs of the two modal encoders pass through the co-attention module for feature crossover; (20) the query input of the co-attention text part transformer is the output of the text encoder, and the key and value are the output of the image encoder; (21) attention is paid through different Q, K and V matrices; (22) the results of attention calculation are added and layer normalization calculated with the input vector of Q, and finally the fused text representation is output; (23) the image part is on the contrary, wherein the query input of the transformer is the output of the image encoder, and the key and value are the output of the text encoder; (24) attention is also calculated through different Q, K and V matrices; and (25) the result of attention calculation is added and layer normalization is calculated with the input vector of Q, and finally the fused image representation is output. Lu also discloses in Section IV with FIG. 5 of Pages 597-599 that (1) since the encoder only obtains the reference from a single modal, it may encounter the problem of modality-bias; (2) therefore, we introduce the extractive text summarization task to supervise the encoder. In other words, we regard extractive text summarization as one of the subtasks; (3) a greedy algorithm similar to Nallapati et al.[25] is used to obtain an oracle summary in each document as an extractive text reference: (4) the extractive summarization is modeled as a text classification task; (5) because the actual data set does not have classification labels, it is necessary to generate pseudo labels first, and use the ROUGE score to select sentences, so that the ROUGE score of the selected sentences and the gold summary is the highest; (6) finally, the generated labels are used to train the model; (7) the abstractive text summarization task adopts the transfer learning method, extracts the text representation by using BART's pre-trained model; (8) after being fused with the image information by the cross attention fusion module, the decoding is completed in the transformer decoder; (9) finally, the abstractive summarization is generated in the abstractive output layer composed of pointer generation network; (10) the text decoding module uses two transformers with multi-head attention mechanism to decode the feature sequence; (11) among them, a multi-head attention takes the text sequence in the picture as input, and finds the association between different characters by learning the context information in the sequence from self-attention; (12) another multi-head attention is used to connect the encoder and decoder to calculate the matching degree between the image features and the text characters; (13) the main advantage of this decoder is that it can learn long-distance dependencies. The structure of the transformer decoder is shown in Figure 5: (14) our image decoder directly uses Multi-Layer Perception, because the text and image have undergone sufficiently complex feature crossover in the co-attention part, and the upper layer only needs to be connected with a simple classifier; and (15) MLP uses two-layer perceptron, the activation layer in the middle uses hyperbolic tangent function, and the final input is converted into probability through sigmoid function.
Curto et al., ("Dyadformer: A Multi-modal Transformer for Long-Range Modeling of Dyadic Interactions", 2021 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), Oct. 11-17, 2021, pp. 177-2188) discloses in Abstract and Section 1 with FIG. 1 of Pages 2177-2178 that (1) present the Dyadformer, a novel multimodal multi-subject Transformer architecture to model individual and interpersonal features in dyadic interactions using variable time windows, thus allowing the capture of long-term interdependencies; (2) our proposed cross-subject layer allows the network to explicitly model interactions among subjects through attentional operations; (3) this proof-of-concept approach shows how multi-modality and joint modeling of both interactants for longer periods of time helps to predict individual attributes; (4) propose a novel architecture to leverage long-term information for joint modeling of both interlocutors in dyadic scenarios; (5) more precisely, we predict the personalities of both interactants by considering not only the audio-visual information and contextual factors (referred to as metadata) independently for each one but also by explicitly modeling their interaction; (6) the proposed model, Dyadformer, mainly consists of two stages: (a) a cross-modal stage where cross-attention encoders fuse multi-modal information, and (v) a cross-subject stage which aims to shape the interaction by performing double cross-attention (see Fig. 1); (7) our contributions are summarized as follows: (a) this method is the first one to jointly model (and infer) self-reported personality in dyadic interactions using time windows of up to ∼30 seconds; (b) inspired by the classical decoder block of the Transformer network [69], we leverage a cross-attention mechanism to both fuse modalities and allow information to flow between subjects; (c) Dyadformer obtained state-of-the-art results on the large-scale UDIVA v0.5 [58] dataset, by reducing previous participant-level error by a 12.5% (from 0.812 to 0.722) when predicting self-reported personality. Curto also discloses in Section 3 of Pages 2179-2181 with FIG. 1 of Page 2177 that (1) present the Dyadformer (Fig. 1), composed of a set of attentional encoder modules, wherein (a) each of these is a stack of Transformer layers [69]; (b) a complete transformer layer is composed of two or more sub-layers; and (c) each sub-layer executes a core block, followed by a residual connection and layer normalization [2]; (2) the core of the Transformer layer is a non-local operation [73], which allows every element in the input sequence to access information from any other; (3) this is achieved through a special form of attention; (4) to compute it, the input representation J is mapped to a set of queries Q, and a memory M is mapped to a set of paired keys K and values V; (5) in the Transformer, the non-local operation is instantiated as the dot-product between Q and K in order to generate an affinity (attention) matrix that weights how much each value should contribute to the augmented representation of every other value; (6) however, in order to build a full transformer layer, attention is not all you need; (7) a complete transformer layer is composed by sub-layers, where the output of one sub-layer is fed as input to the next, n references the attentional encoder module that the sublayer belongs to and Block will either be Multi-Head self-Attention (MHA) or position-wise Feed-Forward Network (FFN) [69] which we define next; (8) in order to allow for the model to attend to different information in a single sub-layer, [69] proposed the Multi-Head self-Attention (MHA) for Sub-Layer blocks; (9) in practice, every sub-layer in a Transformer layer contains a MHA block except for the last one, which contains a FFN block; (10) in our design for Attentional Encoder modules, we use two main modules to build the complete architecture: (a) the self-attention encoder SA(J), which is used to enhance features by attending to themselves, and the cross-attention encoder CA(J,M), which is used to allow for a set of features to attend to a different source; (11) the former is composed of Transformer layers with a single SubLayer, while the latter is composed by two SubLayers; (12) as mentioned earlier, in both cases (SA(J) and CA(J,M)), the described sub-layers are followed by a last SubLayer (i.e., FFN sub-layer); (13) it is important to note that, in the cross-attention encoder, while the input Jl of layer l is the output from the previous cross-encoder layer, M is the same for all layers, which allows Jl features to iteratively attend to M to be progressively augmented; (14) finally, as the self-attention operation is agnostic to relative position among input elements, [69] proposed using positional encodings to indicate the order of the input sequence by a composition of sine and cosine functions at varying frequencies; (15) Dyadformer (depicted in Fig. 1) is a multi-subject multimodal architecture that follows the aforementioned transformer layers; (16) the Dyadformer receives as input a sequence of T small, consecutive and temporally aligned video/audio chunks and infers the personality traits for both subjects in a dyadic interaction; (17) it is composed of two main streams, each of which simultaneously processes a single subject; (18) context and interpersonal features are crucial to predict individual features in dyadic and small group interaction scenarios; (19) for this reason, we propose a model which is capable of (a) fusing information from multiple sources (video, audio, and contextual metadata), and (b) allowing per-subject streams to access each other, in order to consider crossed influence during the interaction; (20) to satisfy both, we go beyond self-attention, where J = M, and also use cross-attention, where J ≠ M; (21) cross-attention works similarly to encoder-decoder attention in [69], where the input and memory come from different sources; (22) for a transformer focusing on dyadic interactions, the target (J) will be from the subject of interest, while the memory (M) will be from the other one; (23) the intuition behind this is to allow information from a given subject to query for useful information from the other; (24) but first, each stream will create an individual representation for each subject; (25) in order to do so, we draw inspiration from multi-modal transformer models [84, 39, 37], and use this same cross-attentional mechanism to fuse data coming from video and audio modalities; (26) in this cross-modal module, J is from the video modality, while M is from the audio one, thus enriching video information with the audio signal; (27) finally, personality scores for both individuals are predicted jointly; (28) in order to build the multi-modal representation for each subject, we first feed the audio features to an audio encoder module composed by layers; then, we use a cross-encoder with layers to enhance video features with the new audio features; (29) the enhanced video features of each subject are transformed through a subject encoder with layers in order to learn rich relationships within individual subject features; and (30) this subject encoder is followed by a cross-encoder with layers as to allow the features from each subject to draw relevant information from each other.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HWEI-MIN LU whose telephone number is (313)446-4913. The examiner can normally be reached Mon - Fri: 9:00 AM - 6:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela D. Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HWEI-MIN LU/Primary Examiner, Art Unit 2142
1 See, for example "MTCA: A Multimodal Summarization Model Based on Two-Stream Cross Attention" to Lu et al., presented in 2022 2nd International Conference on Computer Science, Electronic Information Engineering and Intelligent Control Technology (CEI), Sep 23-25, 2022, 1st paragraph of Section III.A: "ViLBERT (Vision-and-Language BERT)[15] proved that the performance of the two stream model is better than that of the one stream model" .