Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 08/26/2026 has been entered.
Response to Amendment
The previous 35 U.S.C. 112(b) rejections are withdrawn due to Applicant’s amendments.
Response to Arguments
Applicant’s arguments filed 08/26/2026 on pages 9-10 of Remarks regarding the rejection under 35 U.S.C. 102 and 103 with respect to claims 1-20 have been fully considered but are moot. New references Lawhern, Yazdizadeh, Zheng and Louizos have been incorporated below to teach the newly presented limitations.
Claim Objections
Claim 10 is objected to because of the following informalities:
In claim 10, line 2, “fully-connect layer” should read “fully-connected layer”.
Appropriate correction is required.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-6, 8-11, 13-16 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Demir et al. (AutoBayes: Automated Bayesian Graph Exploration for Nuisance-Robust Inference); hereinafter Demir in view of Lawhern et al. (EEGNet: A Compact Convolutional Neural Network for EEG-based Brain-Computer Interfaces); hereinafter Lawhern in view of Yazdizadeh et al. (Ensemble Convolutional Neural Networks for Mode Inference in Smartphone Travel Survey); hereinafter Yazdizadeh in view of Zheng et al. (Rectifying Pseudo Label Learning via Uncertainty Estimation for Domain Adaptive Semantic Segmentation); hereinafter Zheng and in further view of Louizos et al. (THE VARIATIONAL FAIR AUTOENCODER); hereinafter Louizos
Claim 1 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos.
Regarding claim 1, Demir teaches a system for automated construction of an artificial neural network architecture, comprising: (Demir [page 2, paragraph 2]: “In this paper, we propose a systematic automation framework called AutoBayes, which searches for the best inference graph model associated to a Bayesian graph model well-suited to reproduce the training datasets.”)
a set of interfaces and data links configured to receive and send signals, (Demir [page 9, 5 Conclusion and Future Work]: “With the Bayes-Ball algorithm, our method can automatically construct reasonable link connections among classifier, encoder, decoder, nuisance estimator and adversary DNN blocks.”; Note: The link connections are the set of interfaces and data links that transmit signals.)
wherein the signals include datasets of training data, validation data and testing data, (Demir [page 16, A.5 ENSEMBLE LEARNING: STACKED GENERALIZATION]: “We randomly split the data into training set Dtrain and validation set Dtest.”)
wherein the signals include a set of random variable factors in multi-dimensional signals X, and wherein part of the random variable factors are associated with task labels Y to identify, and nuisance variations S; (Demir [page 2, 2 Key Contributions]: “graphical Bayesian models that capture the probabilistic relationship between random variables representing the data features X, task labels Y , nuisance variation labels S”; and [page 18, A.7 DNN Model Parameters]: “When we need to feed S along with 2D data of X into the CNN encoder such as in the model Ds, dimension mismatch poses a problem. We address this issue by using one linear layer to project S into the temporal dimensional space of X and another linear layer to project it into the spatial dimensional space of X.”)
a set of memory banks to store a set of reconfigurable deep neural network (DNN) blocks, (Demir [page 3, 3 AUTOBAYES]: “Depending on the derived conditional independency and pruned factor graphs, DNN blocks for encoder E, decoder D, classifier C, nuisance estimator N and adversary A are reasonably connected.”)
wherein each of the reconfigurable DNN blocks is configured with main task pipeline modules to identify the task labels Y from the multi-dimensional signals X and (Demir [page 2, 2 Key Contributions]: “The ultimate goal is to infer the task label Y from the measured data feature X, which is hindered by the presence of nuisance variations (e.g., inter-subject/session variations) that are (partially) labelled by S.”; and [page 13, Bayesian Graph Model A (Direct Markov)]: “We hence should use a standard classification method, as in Fig. 1(a), to infer Y given X, based on the inference model p(y|x) without involving S and Z.”)
with a set of auxiliary regularization modules to adjust disentanglement between a plurality of latent variables Z and the nuisance variations S, (Demir [page 7, Adversarial Regularization]: “We can utilize adversarial censoring when Z and S should be marginally independent, e.g., such as in Fig. 1(b) and Fig. 7, in order to reinforce the learning of a representation Z that is disentangled from the nuisance variations S. This is accomplished by introducing an adversarial network that aims to maximize a parameterized approximation q(s|z) of the likelihood p(s|z), while this likelihood is also incorporated into the loss for the other modules with a negative weight. The adversarial network, by maximizing the log likelihood log q(sjz), essentially maximizes a lower-bound of the mutual information I(S;Z), and hence the main network is regularized with the additional term that corresponds to minimizing this estimate of mutual information.”)
wherein the memory banks further include hyperparameters, (Demir [page 9]: “AutoBayes can be readily integrated with AutoML to optimize any hyperparameters of individual DNN blocks.”)
trainable variables, (Demir [page 6, Training]: “Each factor block is realized by a DNN, e.g., parameterized by ϴ for
p
ϴ
z
1
z
2
x
”)
intermediate neuron signals, and (Demir [page 18, A.7 DNN Model Parameters]: “When we need to feed S along with 2D data of X into the CNN encoder such as in the model Ds, dimension mismatch poses a problem. We address this issue by using one linear layer to project S into the temporal dimensional space of X and another linear layer to project it into the spatial dimensional space of X.”)
temporary computation values including forward-pass signals and (Demir [page 6, Training]: “Each factor block is realized by a DNN, e.g., parameterized by ϴ for
p
ϴ
z
1
z
2
x
, and all of the networks except for adversarial network are optimized to minimize corresponding loss functions including
L
y
^
,
y
as follows: (6)”; Note: See Equation (6)).
backward-pass gradients; (Demir [page 7, Adversarial Regularization]: “We can utilize adversarial censoring when Z and S should be marginally independent, e.g., such as in Fig. 1(b) and Fig. 7, in order to reinforce the learning of a representation Z that is disentangled from the nuisance variations S. This is accomplished by introducing an adversarial network that aims to maximize a parameterized approximation q(s|z) of the likelihood p(s|z), while this likelihood is also incorporated into the loss for the other modules with a negative weight.”; and [page 4, Algorithm 1]: “Adversary train the whole DNN structure to minimize a loss function”; Note: Minimizing a loss function is done during backward propagation)
at least one processor, in connection with the interface and the memory banks, configured to submit the signals and the datasets into the reconfigurable DNN blocks, (Demir [page 3]: “AutoBayes Algorithm: The overall procedure of the AutoBayes algorithm is described in the pseudocode of Algorithm 1. The AutoBayes automatically constructs non-redundant inference factor graphs given a hypothetical Bayesian graph assumption, through the use of the Bayes-Ball algorithm. Depending on the derived conditional independency and pruned factor graphs, DNN blocks for encoder E, decoder D, classifier C, nuisance estimator N and adversary A are reasonably connected. The whole DNN blocks are trained with adversary learning in a variational Bayesian inference. Note that hyperparameters of each DNN block can be further optimized by AutoML on top of AutoBayes framework.”)
wherein the at least one processor is configured to execute an exploration across combinations of:
a plurality of graphical models, (Demir [page 1, Abstract]: “AutoBayes, that explores different graphical models linking classifier, encoder, decoder, estimator and adversary network blocks to optimize nuisance-invariant machine learning pipelines.”; and [page 4, Algorithm 1] “Attach an adversary network A to latent nodes Z for Zk ⊥ S ∈ I”;)
a plurality of pre-shot regularization methods, (Demir [page 14, Bayesian Graph Model E]: “Note that the generative model E has no marginal dependency between Z and S, which provides the reason to use adversarial censoring to suppress nuisance information S in the latent space Z.”; [page 13, Bayesian Graph Model B]: “Note that this model assumes independence between Z and S, and thus adversarial censoring (Makhzani et al., 2015; Creswell et al., 2017; Lample et al., 2017) can make it more robust against nuisance.”; Note: This is a marginal censoring mode; and [page 14]: “Note that Z2 is marginally independent of the nuisance variable S, which encourages the use of adversarial training to be robust against subject/session variations.”; [page 4]: “Given a Bayesian graph, we can determine whether two disjoint sets of nodes are independent conditionally on other nodes through a graph separation criterion.”; Note: This is a conditional censoring mode)
to reconfigure the reconfigurable DNN blocks such that task prediction is insensitive to the nuisance variations S (Demir [page 7, Adversarial Regularization]: “We can utilize adversarial censoring when Z and S should be marginally independent, e.g., such as in Fig. 1(b) and Fig. 7, in order to reinforce the learning of a representation Z that is disentangled from the nuisance variations S.”)
by modifying the hyperparameters in the memory banks, and (Demir [page 3]: “encoder E, decoder D, classifier C, nuisance estimator N and adversary A are reasonably connected. The whole DNN blocks are trained with adversary learning in a variational Bayesian inference. Note that hyperparameters of each DNN block can be further optimized by AutoML on top of AutoBayes framework.”)
Demir does not appear to explicitly teach a plurality of pre-processing methods,
However, Lawhern teaches a plurality of pre-processing methods, (Lawhern [page 3]: “We introduce the use of Depthwise and Separable convolutions, previously used in computer vision [42], to construct an EEG-specific network that encapsulates several well-known EEG feature extraction concepts, such as optimal spatial filtering and filter-bank construction, while simultaneously reducing the number of trainable parameters to fit when compared to existing approaches.”; [page 7]: “The network starts with a temporal convolution (second column) to learn frequency filters, then uses a depthwise convolution (middle column), connected to each feature map individually, to learn frequency-specific spatial filters.”; and [page 7]: “We apply Batch Normalization [79] along the feature map dimension before applying the exponential linear unit (ELU) nonlinearity [80].” Note: These are spatial filtering, spatio-temporal filtering and normalization used for pre-processing)
It would have been obvious before the effective filing date to combine the hyperparameter optimization of Demir with the spatial filtering used in EEGNet of Lawhern for improved classification (Lawhern, page 3). Demir and Lawhern are analogous art because they both concern deep learning classification of biosignals data.
Demir teaches a [plurality of] post-processing methods, and (Demir [page 7]: “Ensemble Learning: We further introduce ensemble methods to make best use of all Bayesian graph models explored by the AutoBayes framework without wasting lower-performance models. Ensemble stacked generalization works by stacking the predictions of the base learners in a higher level learning space, where a meta learner corrects the predictions of base learners (Wolpert, 1992)”;)
Demir does not appear to explicitly teach a plurality of post-processing methods, and
However, Yazdizadeh teaches a plurality of post-processing methods, and (Yazdizadeh [page 2, C. Ensemble Methods]: "Well-known ensemble techniques include boosting, bagging and stacking. Stacking combines the outputs of a set of base learners and lets another algorithm, referred to as the meta-learner, make the final predictions [14]."; and [page 3]: "The most common ensemble method used for neural networks is average voting that generates posterior labels by calculating the average of the softmax class probabilities or predicted labels for all the base learners [14]."; Note: These teach ensemble stacking and score averaging which are post-processing methods)
It would have been obvious before the effective filing date to combine the hyperparameter optimization of Demir with the ensemble stacking and average voting of Yazdizadeh to improve prediction accuracy of models (Yazdizadeh, page 1, column 2). Demir and Yazdizadeh are analogous art because they both concern classification of signals.
Demir does not appear to explicitly teach a plurality of post-shot adaptation methods,
However, Zheng teaches a plurality of post-shot adaptation methods, (Zheng [page 3, 2.2. Pseudo label learning]: “Another line of semantic segmentation adaptation approaches utilizes the pseudo label to adapt the model to target domain [Zou et al., 2018,Zou et al., 2019,Zheng and Yang, 2020]. The main idea is close to the conventional semi-supervised learning approach, entropy minimization, which is first pro posed to leverage the unlabeled data [Grandvalet and Ben gio, 2005]. Entropy minimization encourages the model to give the prediction with a higher confidence score. In practice, Reed et al. [Reed et al., 2014] propose bootstrapping via entropy minimization and show the effectiveness on the object detection and emotion recognition. Furthermore, Lee et al. [Lee, 2013] exploit the trained model to predict pseudo labels for the unlabeled data, and then fine-tune the model as supervised learning methods to fully leverage the unlabeled data.”; Note: The post-shot adaptation method is the domain adaptation covered in Zheng, examples being pseudo labeling and entropy minimization)
It would have been obvious before the effective filing date to combine the hyperparameter optimization of Demir with the domain adaptation of Zheng to improve pseudo-labeling (Zheng, page 1, column 1). Demir and Zheng are analogous art because they both concern pseudo-labeling in semi-supervised learning.
Demir teaches wherein the hyperparameters are modified to specify the plurality of pre-shot regularization methods using different censoring modes and (Demir [page 14, Bayesian Graph Model E]: “Note that the generative model E has no marginal dependency between Z and S, which provides the reason to use adversarial censoring to suppress nuisance information S in the latent space Z.”; [page 13, Bayesian Graph Model B]: “Note that this model assumes independence between Z and S, and thus adversarial censoring (Makhzani et al., 2015; Creswell et al., 2017; Lample et al., 2017) can make it more robust against nuisance.”; Note: This is a marginal censoring mode; and [page 14]: “Note that Z2 is marginally independent of the nuisance variable S, which encourages the use of adversarial training to be robust against subject/session variations.”; [page 4]: “Given a Bayesian graph, we can determine whether two disjoint sets of nodes are independent conditionally on other nodes through a graph separation criterion.”; Note: This is a conditional censoring mode)
[different] censoring methods. (Demir [page 7, Adversarial Regularization]: “The adversarial network, by maximizing the log likelihood log q(s|z), essentially maximizes a lower-bound of the mutual information I(S;Z), and hence the main network is regularized with the additional term that corresponds to minimizing this estimate of mutual information.”)
Demir does not appear to explicitly teach different censoring methods
However, Louizos teaches different censoring methods (Louizos [page 1, Abstract]: “We investigate the problem of learning representations that are invariant to certain nuisance or sensitive factors of variation in the data while retaining as much of the remaining information as possible. Our model is based on a variational autoencoding architecture (Kingma & Welling, 2014; Rezende et al., 2014) with priors that encourage independence between sensitive and latent factors of variation. Any subsequent processing, such as classification, can then be performed on this purged latent representation. To remove any remaining dependencies we in corporate an additional penalty term based on the “Maximum Mean Discrepancy”)
It would have been obvious before the effective filing date to combine the hyperparameter optimization of Demir with the maximum mean discrepancy of Louizos to encourage independence between sensitive and latent factors of variation (Louizos, page 1, Abstract). Demir and Louizos are analogous art because they both concern learning representations that are invariant to a nuisance factor.
Claim 2 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1. Regarding claim 2, Demir teaches modifying the hyperparameters to specify the plurality of graphical models representing a Bayesian graph model and an inference factor graph using a Bayes-ball algorithm; (Demir [page 3, 3 AutoBayes]: “The AutoBayes automatically constructs non-redundant inference factor graphs given a hypothetical Bayesian graph assumption, through the use of the Bayes-Ball algorithm. Depending on the derived conditional independency and pruned factor graphs, DNN blocks for encoder E, decoder D, classifier C, nuisance estimator N and adversary A are reasonably connected. The whole DNN blocks are trained with adversary learning in a variational Bayesian inference. Note that hyperparameters of each DNN block can be further optimized by AutoML on top of AutoBayes framework.”;)
modifying the reconfigurable DNN blocks by linking graph nodes with graph edges to associate with the random variable factors with respect to the multi-dimensional signals X, the task labels Y, the nuisance variations S and the latent variables Z according to the Bayesian graph model and the inference factor graph; (Demir [page 3]: “AutoBayes offers a solid reason of how to connect multiple DNN blocks to impose conditioning and adversary censoring for the task classifier, feature encoder, decoder, nuisance indicator and adversary networks, based on an explored Bayesian graph”; Note: See Algorithm 1 on page 4 to see dimensional signals X, the task labels Y, the nuisance variations S and the latent variables Z.)
training the reconfigurable DNN blocks with a variational sampling and (Demir [page 3]: “The whole DNN blocks are trained with adversary learning in a variational Bayesian inference. Note that hyperparameters of each DNN block can be further optimized by AutoML on top of AutoBayes framework.”)
a gradient method for the training data; (Demir [page 7, Model Implementation: “All models were trained with a minibatch size of 32 and using the Adam optimizer with an initial learning rate of 0.001. The learning rate is halved whenever the validation loss plateaus.”)
selecting the hyperparameters using an output of the reconfigurable DNN blocks for the validation data; and (Demir [page 4, Algorithm 1]: “return the best model having highest task accuracy in validation sets”)
testing the trained reconfigurable DNN blocks for the testing data and (Demir [page 16, A.5 ENSEMBLE LEARNING: STACKED GENERALIZATION]: “We randomly split the data into training set Dtrain and validation set Dtest … Hold-out Dtest is used to measure the classification performance of both base and meta learners.”)
new incoming data on fly to be transferred with nuisance robustness. (Demir [page 16, QMNIST]: “Additional 297 writers provide 10,000 test samples.”; and [page 8, Results]: “This suggests that we must consider different inference strategies for each target dataset and our AutoBayes provides such an adaptive framework across datasets.”; Note: The additional writers are new and unseen)
Claim 3 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1. Regarding claim 3, Demir teaches Demir teaches wherein the different censoring modes includes at least one of a marginal censoring mode, (Demir [page 14, Bayesian Graph Model E]: “Note that the generative model E has no marginal dependency between Z and S, which provides the reason to use adversarial censoring to suppress nuisance information S in the latent space Z.”; and [page 13, Bayesian Graph Model B]: “Note that this model assumes independence between Z and S, and thus adversarial censoring (Makhzani et al., 2015; Creswell et al., 2017; Lample et al., 2017) can make it more robust against nuisance.”; Note: Adversarial censoring is used to achieve marginal censoring.) a conditional censoring mode, and (Demir [page 14]: “Note that Z2 is marginally independent of the nuisance variable S, which encourages the use of adversarial training to be robust against subject/session variations.”; and [page 4]: “Given a Bayesian graph, we can determine whether two disjoint sets of nodes are independent conditionally on other nodes through a graph separation criterion.”) wherein the different censoring methods includes at least one of (Demir [page 7, Adversarial Regularization]: “The adversarial network, by maximizing the log likelihood log q(s|z), essentially maximizes a lower-bound of the mutual information I(S;Z), and hence the main network is regularized with the additional term that corresponds to minimizing this estimate of mutual information.”)
wherein the at least one processor further executes steps of:
associating the set of auxiliary regularization modules with the reconfigurable DNN blocks such that at least one of latent nodes Z is disentangled from at least one of nuisance variations S according to the set of pre-shot regularization methods; (Demir [page 7, Adversarial Regularization]: “We can utilize adversarial censoring when Z and S should be marginally independent, e.g., such as in Fig. 1(b) and Fig. 7, in order to reinforce the learning of a representation Z that is disentangled from the nuisance variations S.”; Note: See Algorithm 1 of Demir to see the training method in lines 10-15. Adversarial censoring is one of the pre-shot regularization methods.)
training the reconfigurable DNN blocks with the set of auxiliary regularization modules based on the training data; and (Demir [page 3, Algorithm 1, line 16]: “Adversary train the whole DNN structure to minimize a loss function in (5))
selecting the hyperparameters for the different censoring modes and the [different] censoring methods (Demir [page 14, Bayesian Graph Model E]: “Note that the generative model E has no marginal dependency between Z and S, which provides the reason to use adversarial censoring to suppress nuisance information S in the latent space Z.”; [page 13, Bayesian Graph Model B]: “Note that this model assumes independence between Z and S, and thus adversarial censoring (Makhzani et al., 2015; Creswell et al., 2017; Lample et al., 2017) can make it more robust against nuisance.”; Note: This is a marginal censoring mode; and [page 14]: “Note that Z2 is marginally independent of the nuisance variable S, which encourages the use of adversarial training to be robust against subject/session variations.”; [page 4]: “Given a Bayesian graph, we can determine whether two disjoint sets of nodes are independent conditionally on other nodes through a graph separation criterion.”; Note: This is a conditional censoring mode)
[different] censoring methods. (Demir [page 7, Adversarial Regularization]: “The adversarial network, by maximizing the log likelihood log q(s|z), essentially maximizes a lower-bound of the mutual information I(S;Z), and hence the main network is regularized with the additional term that corresponds to minimizing this estimate of mutual information.”)
Demir does not appear to explicitly teach different censoring methods
However, Louizos teaches different censoring methods (Louizos [page 1, Abstract]: “We investigate the problem of learning representations that are invariant to certain nuisance or sensitive factors of variation in the data while retaining as much of the remaining information as possible. Our model is based on a variational autoencoding architecture (Kingma & Welling, 2014; Rezende et al., 2014) with priors that encourage independence between sensitive and latent factors of variation. Any subsequent processing, such as classification, can then be performed on this purged latent representation. To remove any remaining dependencies we in corporate an additional penalty term based on the “Maximum Mean Discrepancy”)
It would have been obvious before the effective filing date to combine the hyperparameter optimization of Demir with the maximum mean discrepancy of Louizos to encourage independence between sensitive and latent factors of variation (Louizos, page 1, Abstract). Demir and Louizos are analogous art because they both concern learning representations that are invariant to a nuisance factor.
Claim 4 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1.
Regarding claim 4, Demir teaches wherein the different censoring methods include at least one of adversarial censoring method, a mutual information neural estimation (MINE) censoring method, a mutual information gradient estimation (MIGE) censoring method, a maximum mean discrepancy (MMD) censoring method, a pairwise maximum mean discrepancy (MMD) censoring method, a boundary equilibrium generative adversarial network (BEGAN) discriminator censoring method, a Hilbert-Schmidt independence criterion (HSIC) censoring method, and an optimal transport censoring method. (Demir [page 14, Bayesian Graph Model E]: “Note that the generative model E has no marginal dependency between Z and S, which provides the reason to use adversarial censoring to suppress nuisance information S in the latent space Z.”; and [page 13, Bayesian Graph Model B]: “Note that this model assumes independence between Z and S, and thus adversarial censoring (Makhzani et al., 2015; Creswell et al., 2017; Lample et al., 2017) can make it more robust against nuisance.”; Note: Adversarial censoring is used to achieve marginal censoring.)
Claim 5 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1.
Regarding claim 5, Demir teaches modifying the training data, the validation data, and the testing data to feed in the reconfigurable DNN blocks [according to the plurality of pre-processing methods.] (See Algorithm 1 of Demir to see lines 1-20)
Demir does not appear to explicitly teach modifying the hyperparameters to specify the plurality of pre-processing methods, modifying the hyperparameters includes using at least one of a spatial filtering, spatio-temporal filtering, wavelet transforms, vector auto-regressive filter, self-attention mapping, robust z-scoring, normalization, data augmentation, and universal adversarial example; and
according to the plurality of pre-processing methods.
However, Lawhern teaches modifying the hyperparameters to specify the plurality of pre-processing methods, modifying the hyperparameters includes using at least one of a spatial filtering, spatio-temporal filtering, wavelet transforms, vector auto-regressive filter, self-attention mapping, robust z-scoring, normalization, data augmentation, and universal adversarial example; and
according to the plurality of pre-processing methods. (Lawhern [page 3]: “We introduce the use of Depthwise and Separable convolutions, previously used in computer vision [42], to construct an EEG-specific network that encapsulates several well-known EEG feature extraction concepts, such as optimal spatial filtering and filter-bank construction, while simultaneously reducing the number of trainable parameters to fit when compared to existing approaches.”; [page 7]: “The network starts with a temporal convolution (second column) to learn frequency filters, then uses a depthwise convolution (middle column), connected to each feature map individually, to learn frequency-specific spatial filters.”; and [page 7]: “We apply Batch Normalization [79] along the feature map dimension before applying the exponential linear unit (ELU) nonlinearity [80].” Note: These are spatial filtering, spatio-temporal filtering and normalization used for pre-processing)
It would have been obvious before the effective filing date to combine the hyperparameter optimization of Demir with the spatial filtering used in EEGNet of Lawhern for improved classification (Lawhern, page 3). Demir and Lawhern are analogous art because they both concern deep learning classification of biosignals data.
Claim 6 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1.
Regarding claim 6, Demir teaches wherein the plurality of post-processing methods includes at least one of (Demir [page 7]: “Ensemble Learning: We further introduce ensemble methods to make best use of all Bayesian graph models explored by the AutoBayes framework without wasting lower-performance models. Ensemble stacked generalization works by stacking the predictions of the base learners in a higher level learning space, where a meta learner corrects the predictions of base learners (Wolpert, 1992)”;)
Claim 8 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1.
Demir teaches wherein the variational sampling is employed for the latent variables with an independent distribution specified by an exponential family or non-exponential family, as its prior distribution for reparameterization tricks, and for categorical variables of unknown nuisance variations and task labels using Gumbel softmax trick to produce near-one-hot vectors based on a random number generator and a softmax temperature. (Demir [pages 15-16; Variational Categorical Reparameterization]: “In order to deal with the issue of categorical sampling, we can use the Gumbel-Softmax reparameterization trick (Jang et al., 2016), which enables differentiable approximation of one-hot encoding. Let [π1, π2, . . . , π|S|] denote a target probability mass function for the categorical variable S. Let g1, g2, . . . , g|S| be independent and identically distributed samples drawn from the Gumbel distribution Gumbel(0, 1). 1 Then, generate an |S|-dimensional vector sˆ = [ˆs1, sˆ2, . . . , sˆ|S|] according to (20) where τ > 0 is a softmax temperature. As the softmax temperature τ approaches 0, samples from the Gumbel-Softmax distribution become one-hot and the distribution becomes identical to the target categorical distribution. The temperature τ is usually decreased across training epochs as an annealing technique, e.g., with exponential decaying.”; and [page 15, Graphical Models for Semi-Supervised Learning]: “Nuisance values S such as subject ID or session ID may not be always available for typical physiological datasets, in particular for the testing phase of an HMI system deployment with new users, requiring semi-supervised methods. We note that some graphical models are well-suited for such semi-supervised training.”; Note: g is a random variable and see page 16 to see that task label |Y| is a categorical variable processed throughout Demir)
Claim 9 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1.
Regarding claim 9, Demir teaches wherein link concatenation comprises a step of multi-dimensional tensor projection with a plurality of trainable linear filters or bilinear filters to convert lower-dimensional signals for dimension-mismatched links. (Demir [page 18, A.7 DNN Model Parameters]: “When we need to feed S along with 2D data of X into the CNN encoder such as in the model Ds, dimension mismatch poses a problem. We address this issue by using one linear layer to project S into the temporal dimensional space of X and another linear layer to project it into the spatial dimensional space of X. The dot product of those two projected vectors is concatenated as additional channel input.”;)
Claim 10 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1.
Regarding claim 10, Demir teaches wherein the reconfigurable DNN blocks are configured with a combination of at least two fully-connect layer, convolutional layer, graph convolutional layer, recurrent layer, loopy connection, skip connection, and inception layer with a set of nonlinear activations including at least one of rectified linear variants, hyperbolic tangent, sigmoid, gated linear, softmax, and thresholding, regularized with a combination of dropout, swap out, zone out, block out, drop connect, noise injection, shaking, and batch normalization. (Demir [page 18, A.7 DNN Model Parameters, Table 3]: “For 2D datasets, we use deep CNN for the encoder E and decoder D blocks. For the classifier C, nuisance estimator N, and adversary A, we use a multi-layer perceptron (MLP) having three layers, whose hidden nodes are doubled from the input dimension. We also use batch normalization (BN) and ReLU activation as listed in Table 3.”)
Claim 11 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1.
Regarding claim 11, Demir teaches wherein the training performs updating the trainable parameters of the reconfigurable DNN blocks by using the training data such that output of the reconfigurable DNN blocks provide smaller loss values in a combination of at least two of objective functions, wherein the objective functions further include a combination of mean-square error, cross entropy, structural similarity, negative log-likelihood, absolute error, cross covariance, clustering loss, divergence, hinge loss, Huber loss, negative sampling, Wasserstein distance, and triplet loss, wherein the loss functions are weighted with a plural of regularization coefficients adjusted according to the specified training schedules. (See Equation 7 of Demir on page 7 to see the combination of Kullback–Leibler divergence and negative log-likelihood)
Claim 13 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1. Regarding claim 13, Demir teaches wherein the datasets includes sensor measurements further comprising one or more of:
media data including images, pictures, movies, texts, letters, voices, music, audios, speeches; (Demir [page 16, A.6 Datasets Description]: “QMNIST: A hand-written digit image MNIST with extended label information including a writer ID number”;)
physical data including radio waves, optical signals, electrical pulses, temperatures, pressures, accelerations, speeds, vibrations and forces; and (Demir [page 16, A.6 Datasets Description]: “The data were collected by C = 7 sensors, i.e., electrodermal activity, temperature, three-dimensional acceleration, heart rate, and arterial oxygen level.”;)
physiological data including heart rate, blood pressure, mass, moisture, electroencephalogram, electromyogram, electrocardiogram, mechanomyogram, electrooculogram, galvanic skin response, magnetoencephalogram, and electrocorticography. (Demir [page 17, A.6 Datasets Description]: “The dataset consists of EEG data recorded from |S| = 16 healthy subjects participating in an offline P300 spelling task, where visual feedback of the inferred letter is provided to the user at the end of each trial for 1.3 seconds to monitor evoked brain responses for erroneous decisions made by the system.”;)
Claim 14 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1. Regarding claim 14, Demir teaches wherein the nuisance variations include a set of subject identifications, (Demir [page 2]: “In order to promote robustness against nuisance parameters such as subject IDs, the explored Bayesian graphs can provide reasoning to use adversarial training with/without variational modeling and latent disentanglement.”)
session numbers, (Demir [page 2]: “The ultimate goal is to infer the task label Y from the measured data feature X, which is hindered by the presence of nuisance variations (e.g., inter-subject/session variations) that are (partially) labelled by S.”)
Claim 15 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1.
Regarding claim 15, Demir teaches wherein each of the reconfigurable DNN block further comprises hyperparameters specifying a set of layers having a set of artificial neuron nodes, (Demir [page 18, A.7 DNN Model Parameters, Table 3]: “For 2D datasets, we use deep CNN for the encoder E and decoder D blocks. For the classifier C, nuisance estimator N , and adversary A, we use a multi-layer perceptron (MLP) having three layers, whose hidden nodes are doubled from the input dimension. We also use batch normalization (BN) and ReLU activation as listed in Table 3.”) wherein a pair of the neuron nodes from neighboring layers are mutually connected with a plural of trainable variables and (Demir [page 6, Training]: “Each factor block is realized by a DNN, e.g., parameterized by ϴ for
p
ϴ
z
1
z
2
x
). activation functions to pass a signal from the previous layers to the next layers sequentially. (Demir [page 8, Model Implementation]: “Each convolution is followed by a batch normalization (BN) and a rectified linear unit (ReLU) activation.”)
Claim 16 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1.
Regarding claim 16, Demir teaches wherein the nuisance variations S are further decomposed into multiple factors of variations S1, S2, . . . , SN as multiple-domain side information according to at least one of supervised, semi-supervised and 1, Z2, . . . , ZL as disentangled feature vectors. (Demir [page 4, Algorithm 1]: “Require: Nodes set V = [Y, X, S1, S2, … , Sn, Z1, Z2, … , Zm], where Y denotes task labels, X is a measurement data, S = [S1, S2, … , Sn] are (potentially multiple) semi-supervised nuisance variations, and Z = [Z1, Z2, … , Zm] are (potentially multiple) latent vectors.)
Claim 18 is rejected over Demir, Lawhern, Yazdizadeh, Zheng and Louizos with the incorporation of claim 1.
Regarding claim 18, Demir teaches wherein the hyperparameters comprise a set of training schedules including an adaptive control of learning rates, regularization weights, factorization permutations, and policy to prune less-priority links, by using a belief propagation to measure a discrepancy between the training data and the validation data. (Demir [page 7, Model Implementation]: “Model Implementation: All models were trained with a minibatch size of 32 and using the Adam optimizer with an initial learning rate of 0.001. The learning rate is halved whenever the validation loss plateaus.”; Note: This is showing the adaptive control of learning rates)
Claim 19 is claim 1 in the form of a method and is rejected for the same reasons as claim 1 stated above.
Dependent claim 20 is claims 3 and 4 in the form of a method and is rejected for the same reasons as claims 3 and 4 stated above. For the rejection of the limitations specifically pertaining to the method of claim 19, see the rejection of claim 19 above.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Demir, Lawhern, Yazdizadeh, Zheng and Louizos and in further view of Howard et al. (Universal Language Model Fine-tuning for Text Classification); hereinafter Howard
Claim 7 is rejected over Demir, Lawhern, Yazdizadeh, Zheng, Louizos and Howard with the incorporation of claim 1. Regarding claim 7, Demir does not appear to explicitly teach wherein the plurality of post-shot adaptation methods includes at least one of pseudo-labeling, soft labeling, confusion minimization, entropy minimization, feature normalization, weighted z-scoring, elastic weight consolidation, label propagation, adaptive layer freezing, hyper network adaptation, latent space clustering, quantization, and sparsification,
However, Zheng teaches wherein the plurality of post-shot adaptation methods includes at least one of pseudo-labeling, soft labeling, confusion minimization, entropy minimization, feature normalization, weighted z-scoring, elastic weight consolidation, label propagation, adaptive layer freezing, hyper network adaptation, latent space clustering, quantization, and sparsification, (Zheng [page 3, 2.2. Pseudo label learning]: “Another line of semantic segmentation adaptation approaches utilizes the pseudo label to adapt the model to target domain [Zou et al., 2018,Zou et al., 2019,Zheng and Yang, 2020]. The main idea is close to the conventional semi-supervised learning approach, entropy minimization, which is first pro posed to leverage the unlabeled data [Grandvalet and Ben gio, 2005]. Entropy minimization encourages the model to give the prediction with a higher confidence score. In practice, Reed et al. [Reed et al., 2014] propose bootstrapping via entropy minimization and show the effectiveness on the object detection and emotion recognition. Furthermore, Lee et al. [Lee, 2013] exploit the trained model to predict pseudo labels for the unlabeled data, and then fine-tune the model as supervised learning methods to fully leverage the unlabeled data.”; Note: The post-shot adaptation method is the domain adaptation covered in Zheng)
It would have been obvious before the effective filing date to combine the hyperparameter optimization of Demir with the domain adaptation of Zheng to improve pseudo-labeling (Zheng, page 1, column 1). Demir and Zheng are analogous art because they both concern pseudo-labeling in semi-supervised learning.
Demir does not appear to explicitly teach wherein the reconfigurable DNN blocks are refined by unfreezing a combination of the trainable variables such that the reconfigurable DNN blocks adapt to a new-domain dataset.
However, Howard teaches wherein the reconfigurable DNN blocks are refined by unfreezing a combination of the trainable variables such that the reconfigurable DNN blocks adapt to a new-domain dataset. (Howard [page 3, Figure 1]: “The classifier is fine-tuned on the target task using gradual unfreezing, ‘Discr’, and STLR to preserve low-level representations and adapt high-level ones (shaded: unfreezing stages; black: frozen).”)
It would have been obvious before the effective filing date to combine the hyperparameter optimization of Demir with the gradual unfreezing fine-tuning of Howard to improve transfer learning fine-tuning (Howard, page 1, Abstract). Demir and Howard are analogous art because they both concern transfer learning and domain adaptation for deep neural networks.
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Demir, Lawhern, Yazdizadeh, Zheng and Louizos in view of Kingma et al. (ADAM: A METHOD FOR STOCHASTIC OPTIMIZATION); hereinafter Kingma
Claim 12 is rejected over Demir, Lawhern, Yazdizadeh, Zheng, Louizos and Kingma with the incorporation of claim 1.
Regarding claim 12, Demir teaches wherein the gradient method employs a [combination of at least two] of stochastic gradient descent, adaptive momentum, Ada gradient, Ada bound, Nesterov accelerated gradient, and root-mean-square propagation for optimizing trainable parameters of the reconfigurable DNN blocks. (Demir [page 7, Model Implementation]: “Model Implementation: All models were trained with a minibatch size of 32 and using the Adam optimizer with an initial learning rate of 0.001. The learning rate is halved whenever the validation loss plateaus.”; Note: Adam is the adaptive momentum)
Demir does not appear to explicitly teach wherein the gradient method employs a combination of at least two of stochastic gradient descent, adaptive momentum, Ada gradient, Ada bound, Nesterov accelerated gradient, and root-mean-square propagation for optimizing trainable parameters of the reconfigurable DNN blocks
However, Kingma teaches wherein the gradient method employs a combination of at least two of stochastic gradient descent, adaptive momentum, Ada gradient, Ada bound, Nesterov accelerated gradient, and root-mean-square propagation for optimizing trainable parameters of the reconfigurable DNN blocks (Kingma [page 1, 1 INTRODUCTION]: “Our method is designed to combine the advantage of two recently popular methods: AdaGrad (Duchi et al., 2011), which works well with sparse gradients, and RMSProp (Tieleman & Hinton, 2012), which works well in on-line and non-stationary settings; important connections to these and other stochastic optimization methods are clarified in section 5.”; Note: AdaGrad = Ada gradient, RMSProp = root-mean-square propagation)
It would have been obvious before the effective filing date to combine the adaptive momentum Adam of Demir with the AdaGrad and RMSProp of Kingma for efficient stochastic optimization with little memory requirement (Kingma, page 1, 1 INTRODUCTION). Demir and Kingma are analogous art because they both concern adaptive momentum.
Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Demir, Lawhern, Yazdizadeh, Zheng and Louizos in view of Falkner et al. (BOHB: Robust and Efficient Hyperparameter Optimization at Scale); hereinafter Falkner
Claim 17 is rejected over Demir, Lawhern, Yazdizadeh, Zheng, Louizos and Falkner with the incorporation of claim 1.
Regarding claim 17, Demir does not appear to explicitly teach wherein the hyperparameters are modified employing at least one of reinforcement learning, evolutionary strategy, differential evolution, particle swarm, genetic algorithm, annealing, Bayesian optimization, hyperband, and multi-objective Lamarckian evolution, to explore different combinations of discrete and continuous hyperparameter values.
However, Falkner teaches wherein the hyperparameters are modified employing at least one of reinforcement learning, evolutionary strategy, differential evolution, particle swarm, genetic algorithm, annealing, Bayesian optimization, hyperband, and multi-objective Lamarckian evolution, to explore different combinations of discrete and continuous hyperparameter values. (Falkner [page 1, Abstract]: “we propose to combine the benefits of both Bayesian optimization and bandit-based methods, in order to achieve the best of both worlds: strong anytime performance and fast convergence to optimal configurations.”; Note: Hyperband is a bandit-based method)
It would have been obvious before the effective filing date to combine the hyperparameter optimization of Demir with the Bayesian optimization and bandit-based methods of Falkner for strong performance and fast convergence to optimal configurations (Falkner, page 1, Abstract). Demir and Falkner are analogous art because they both concern hyperparameter optimization.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID H TRAN whose telephone number is (703)756-1525. The examiner can normally be reached M-F 9:30 am - 5:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DAVID H TRAN/Examiner, Art Unit 2147
/ERIC NILSSON/Primary Examiner, Art Unit 2151