DETAILED ACTION
1. The present application is being examined under the pre-AIA first to invent provisions.
Priority
2. Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Response to Amendment
3. Receipt of Applicant’s Amendment filed on 08/31/2026 is acknowledged. The amendment includes the amending of claims 1-2, 5, 11-13, and 15.
Claim Rejections - 35 USC § 112
4. The rejections raised in the Office Action mailed on 05/29/2026 have been overcome by applicant’s amendment received on 08/31/2026.
Claim Rejections - 35 USC § 103
5. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
6. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
7. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
8. Claims 1 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Jia et al. (U.S. PGPUB 2023/0139567), in view of Kim et al. (U.S. PGPUB 2021/0398004).
9. Regarding claim 1, Jia teaches a method comprising:
A) training, by a multimodal unsupervised feature representation learning unit, an encoder configured to extract features of individual single-modal signals from a source multimodal dataset (Paragraphs 48 and 54);
B) generating, by a multimodal unsupervised task generation unit, a source task based on the features of individual single-modal signals (Paragraphs 48 and 54-55);
C) deriving, by a multimodal unsupervised learning method derivation unit, a learning method from the source task using the encoder (Paragraphs 48 and 54-55); and
D) training, by a target task performance unit, a model based on the learning method and features extracted from target datasets by the encoder, thus performing a target task (Paragraphs 48 and 54-55).
The examiner notes that Jia teaches “training, by a multimodal unsupervised feature representation learning unit, an encoder configured to extract features of individual single-modal signals from a source multimodal dataset” as “A VAE or other autoencoder, such as a MMAE or MVAE, may be trained with inputs from two or more types of multi-omics or multi-modal data. The inputs may include combinations of genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data. In some examples, the input data is genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, and/or phenomic, which includes but is not limited to genomic-wide data, exomic-wide data, epigenomic-wide data, transcriptomic-wide data, proteomic-wide data, metabolomic-wide data, hyperspectral data, and/or phenomic-data or combinations thereof” (Paragraph 48) and “In one embodiment, in the training stage, the multi-modal autoencoder, which includes an encoder trained to encode the input data (the features from the input data) obtained from the training and testing populations into a latent representation in the latent space and a decoder to decode the latent representation and to reconstruct the input data using unsupervised learning. In some aspects, the decoder is trained on existing genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data, or combinations thereof. The training of the autoencoder to learn the encoding and decoding of the two or more types of data or multi-omics data may be iterative until the likelihood of reconstructing the input data reaches a certain level or threshold of accuracy” (Paragraph 54). The examiner further notes that an autoencoder is trained via features of multiple modals of source multi-modal data. The examiner further notes that Jia teaches “generating, by a multimodal unsupervised task generation unit, a source task based on the features of individual single-modal signals” as “A VAE or other autoencoder, such as a MMAE or MVAE, may be trained with inputs from two or more types of multi-omics or multi-modal data. The inputs may include combinations of genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data. In some examples, the input data is genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, and/or phenomic, which includes but is not limited to genomic-wide data, exomic-wide data, epigenomic-wide data, transcriptomic-wide data, proteomic-wide data, metabolomic-wide data, hyperspectral data, and/or phenomic-data or combinations thereof” (Paragraph 48), “In one embodiment, in the training stage, the multi-modal autoencoder, which includes an encoder trained to encode the input data (the features from the input data) obtained from the training and testing populations into a latent representation in the latent space and a decoder to decode the latent representation and to reconstruct the input data using unsupervised learning. In some aspects, the decoder is trained on existing genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data, or combinations thereof. The training of the autoencoder to learn the encoding and decoding of the two or more types of data or multi-omics data may be iterative until the likelihood of reconstructing the input data reaches a certain level or threshold of accuracy” (Paragraph 54), and “a neural network or an autoencoder with an encoder and decoder is trained using unsupervised learning or any other suitable technique. Unsupervised learning is a method that may be used for training the autoencoder in a state in which a label is not allocated to training data. Unsupervised learning includes the task of trying to find hidden structure in unlabeled data. Some examples of unsupervised learning processes include but are not limited to: clustering (e.g., k-means, mixture models, hierarchical clustering), blind signal separation using feature extraction techniques for dimensionality reduction (e.g., principal component analysis, independent component analysis, non-negative matrix factorization, singular value decomposition) and artificial neural networks (e.g., self-organizing map, adaptive resonance theory). Clustering analysis is the assignment of a set of observations into subsets (called clusters) so that observations within the same cluster are similar according to some pre-designated criterion or criteria, while observations drawn from different clusters are dissimilar. Different clustering techniques make different assumptions on the structure of the data, often defined by some similarity metric and evaluated for example by internal compactness (similarity between members of the same cluster) and separation between different clusters. An unsupervised algorithm, e.g., a clustering or dimensionality reduction algorithm, may find previously unknown patterns in data sets without pre-existing labels. Accordingly, in at least one embodiment, the untrained autoencoder trains itself using unlabeled data. In some aspects, the unsupervised learning training dataset includes two or more types of input data, such as multi-omics or multi-modal data, without any associated output data. The training data set may be obtained from a training population, a testing population, or both. The autoencoder through training using unsupervised learning becomes capable of finding hidden structure in unlabeled data, learning groupings within training dataset, determining how individual inputs are related to the untrained dataset, or combinations thereof” (Paragraph 55). The examiner further notes that the iterative training of an autoencoder via unsupervised learning of features of multi-modal data entails the generation of the undefined source task in the broadest reasonable interpretation. Moreover, the instant specification simply defines a “task” as “an issue intended to address through machine learning or a task intended to perform through machine-learning” (Paragraph 42). The examiner further notes that Jia teaches “deriving, by a multimodal unsupervised learning method derivation unit, a learning method from the source task using the encoder” as “A VAE or other autoencoder, such as a MMAE or MVAE, may be trained with inputs from two or more types of multi-omics or multi-modal data. The inputs may include combinations of genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data. In some examples, the input data is genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, and/or phenomic, which includes but is not limited to genomic-wide data, exomic-wide data, epigenomic-wide data, transcriptomic-wide data, proteomic-wide data, metabolomic-wide data, hyperspectral data, and/or phenomic-data or combinations thereof” (Paragraph 48), “In one embodiment, in the training stage, the multi-modal autoencoder, which includes an encoder trained to encode the input data (the features from the input data) obtained from the training and testing populations into a latent representation in the latent space and a decoder to decode the latent representation and to reconstruct the input data using unsupervised learning. In some aspects, the decoder is trained on existing genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data, or combinations thereof. The training of the autoencoder to learn the encoding and decoding of the two or more types of data or multi-omics data may be iterative until the likelihood of reconstructing the input data reaches a certain level or threshold of accuracy” (Paragraph 54), and “a neural network or an autoencoder with an encoder and decoder is trained using unsupervised learning or any other suitable technique. Unsupervised learning is a method that may be used for training the autoencoder in a state in which a label is not allocated to training data. Unsupervised learning includes the task of trying to find hidden structure in unlabeled data. Some examples of unsupervised learning processes include but are not limited to: clustering (e.g., k-means, mixture models, hierarchical clustering), blind signal separation using feature extraction techniques for dimensionality reduction (e.g., principal component analysis, independent component analysis, non-negative matrix factorization, singular value decomposition) and artificial neural networks (e.g., self-organizing map, adaptive resonance theory). Clustering analysis is the assignment of a set of observations into subsets (called clusters) so that observations within the same cluster are similar according to some pre-designated criterion or criteria, while observations drawn from different clusters are dissimilar. Different clustering techniques make different assumptions on the structure of the data, often defined by some similarity metric and evaluated for example by internal compactness (similarity between members of the same cluster) and separation between different clusters. An unsupervised algorithm, e.g., a clustering or dimensionality reduction algorithm, may find previously unknown patterns in data sets without pre-existing labels. Accordingly, in at least one embodiment, the untrained autoencoder trains itself using unlabeled data. In some aspects, the unsupervised learning training dataset includes two or more types of input data, such as multi-omics or multi-modal data, without any associated output data. The training data set may be obtained from a training population, a testing population, or both. The autoencoder through training using unsupervised learning becomes capable of finding hidden structure in unlabeled data, learning groupings within training dataset, determining how individual inputs are related to the untrained dataset, or combinations thereof” (Paragraph 55). The examiner further notes that the iterative training of an autoencoder via unsupervised learning of features of multi-modal data entails the use of a derived “learning method” (which is undefined in the claims) for training/learning the autoencoder in and of itself. The examiner further notes that Jia teaches “training, by a target task performance unit, a model based on the learning method and features extracted from target datasets by the encoder, thus performing a target task” as “A VAE or other autoencoder, such as a MMAE or MVAE, may be trained with inputs from two or more types of multi-omics or multi-modal data. The inputs may include combinations of genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data. In some examples, the input data is genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, and/or phenomic, which includes but is not limited to genomic-wide data, exomic-wide data, epigenomic-wide data, transcriptomic-wide data, proteomic-wide data, metabolomic-wide data, hyperspectral data, and/or phenomic-data or combinations thereof” (Paragraph 48), “In one embodiment, in the training stage, the multi-modal autoencoder, which includes an encoder trained to encode the input data (the features from the input data) obtained from the training and testing populations into a latent representation in the latent space and a decoder to decode the latent representation and to reconstruct the input data using unsupervised learning. In some aspects, the decoder is trained on existing genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data, or combinations thereof. The training of the autoencoder to learn the encoding and decoding of the two or more types of data or multi-omics data may be iterative until the likelihood of reconstructing the input data reaches a certain level or threshold of accuracy” (Paragraph 54), and “a neural network or an autoencoder with an encoder and decoder is trained using unsupervised learning or any other suitable technique. Unsupervised learning is a method that may be used for training the autoencoder in a state in which a label is not allocated to training data. Unsupervised learning includes the task of trying to find hidden structure in unlabeled data. Some examples of unsupervised learning processes include but are not limited to: clustering (e.g., k-means, mixture models, hierarchical clustering), blind signal separation using feature extraction techniques for dimensionality reduction (e.g., principal component analysis, independent component analysis, non-negative matrix factorization, singular value decomposition) and artificial neural networks (e.g., self-organizing map, adaptive resonance theory). Clustering analysis is the assignment of a set of observations into subsets (called clusters) so that observations within the same cluster are similar according to some pre-designated criterion or criteria, while observations drawn from different clusters are dissimilar. Different clustering techniques make different assumptions on the structure of the data, often defined by some similarity metric and evaluated for example by internal compactness (similarity between members of the same cluster) and separation between different clusters. An unsupervised algorithm, e.g., a clustering or dimensionality reduction algorithm, may find previously unknown patterns in data sets without pre-existing labels. Accordingly, in at least one embodiment, the untrained autoencoder trains itself using unlabeled data. In some aspects, the unsupervised learning training dataset includes two or more types of input data, such as multi-omics or multi-modal data, without any associated output data. The training data set may be obtained from a training population, a testing population, or both. The autoencoder through training using unsupervised learning becomes capable of finding hidden structure in unlabeled data, learning groupings within training dataset, determining how individual inputs are related to the untrained dataset, or combinations thereof” (Paragraph 55). The examiner further notes that the iterative training of an autoencoder via unsupervised learning of features of multi-modal data entails the performing of the undefined target task in the broadest reasonable interpretation. Moreover, the instant specification simply defines a “task” as “an issue intended to address through machine learning or a task intended to perform through machine-learning” (Paragraph 42).
Jia does not explicitly teach:
E) wherein deriving the learning method comprises: inputting, by a task estimator, the individual single-model signals to a neural network to estimate a task intended to be performed; and
F) modulating, by a task encoding modulator, the encoder based on the estimated task.
Kim, however, teaches “wherein deriving the learning method comprises: inputting, by a task estimator, the individual single-model signals to a neural network to estimate a task intended to be performed” as “The domain and task estimator 115 estimates domains and tasks of all pieces of the input support data D.sup.k′,t based on the embedded feature information according to the embedding result. In one embodiment, the domain and task estimator 115 may set the embedded feature information as an input of a multi-layer perceptron model and acquire the estimated domain and task of the support data as the output corresponding to the input. In this case, a dimension of an output stage for an output of the multi-layer perceptron model may be set to be smaller than that of an input stage for input” (Paragraphs 65-66) and “modulating, by a task encoding modulator, the encoder based on the estimated task” as “The modulation information acquirer 120 acquires the modulation information of the initial parameter {tilde over (θ)}.sup.k of the task execution model based on the estimated domain and task. In one embodiment, the modulation information acquirer 120 may acquire the modulation information of the initial parameter {tilde over (θ)}.sup.k of the task execution model from the knowledge memory 130 using the estimated domain and task directly from the estimated domain and task or through a knowledge controller 125” (Paragraphs 67-68) and “The modulator 135 modulates the initial parameter {tilde over (θ)}.sup.k of the task execution model based on the modulation information. In this case, the modulator 135 may sum the modulation information directly acquired by the modulation information acquirer 120 and the modulation information acquired from the knowledge memory 130 by the knowledge controller 125 and may modulate the initial parameter {tilde over (θ)}.sup.k of the task execution model based on the summed modulation information” (Paragraph 73).
The examiner further notes that Kim teaches the estimation of a task to be performed via the input of support data to a model. The combination would result in inputting the single-model signals of Jia to estimate a task. Moreover, Kim teaches the modulation of parameters based on the estimated task. The combination would result in the modulation of the autoencoder of Jia via such modulated parameters.
It would have been obvious to one of ordinary skill in the art before the effective filing date of instant invention to combine the teachings of the cited references because teaching Kim’s would have allowed Jia’s to provide a method for coping with diverse domains and tasks, as noted by Kim (Paragraph 57).
Regarding claim 11, Jia teaches an apparatus comprising:
A) a memory configured to store a control program for multimodal unsupervised meta-learning (Paragraphs 22 and 54); and
B) a processor configured to execute the control program stored in the memory (Paragraph 22);
C) wherein the processor is configured to train an encoder configured to extract features of individual single-modal signals from a source multimodal dataset (Paragraphs 48 and 54);
D) generate a source task based on the features of individual single-modal signals (Paragraphs 48 and 54-55);
E) derive a learning method from the source task using the encoder (Paragraphs 48 and 54-55); and
F) perform a target task by training a model based on the features extracted by the encoder and the learning method (Paragraphs 48 and 54-55).
The examiner notes that Jia teaches “a memory configured to store a control program for multimodal unsupervised meta-learning” as “The computing device 110 includes a processor 112, a memory 114, an input/output (I/O) controller 116 (e.g., a network transceiver), a memory unit 118, and a database 120, all of which may be interconnected via one or more address/data bus” (Paragraph 22) and “in the training stage, the multi-modal autoencoder, which includes an encoder trained to encode the input data (the features from the input data) obtained from the training and testing populations into a latent representation in the latent space and a decoder to decode the latent representation and to reconstruct the input data using unsupervised learning” (Paragraph 54). The examiner further notes that the memory 114 teaches the claimed memory. Moreover, the training (i.e. learning) via unsupervised learning of an multi-modal autoencoder teaches the claimed multimodal unsupervised meta-learning in the broadest reasonable interpretation. The examiner further notes that Jia teaches “a processor configured to execute the control program stored in the memory” as “The computing device 110 includes a processor 112, a memory 114, an input/output (I/O) controller 116 (e.g., a network transceiver), a memory unit 118, and a database 120, all of which may be interconnected via one or more address/data bus” (Paragraph 22). The examiner further notes that processor 112 teaches the claimed processor. The examiner further notes that Jia teaches “wherein the processor is configured to train an encoder configured to extract features of individual single-modal signals from a source multimodal dataset” as “A VAE or other autoencoder, such as a MMAE or MVAE, may be trained with inputs from two or more types of multi-omics or multi-modal data. The inputs may include combinations of genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data. In some examples, the input data is genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, and/or phenomic, which includes but is not limited to genomic-wide data, exomic-wide data, epigenomic-wide data, transcriptomic-wide data, proteomic-wide data, metabolomic-wide data, hyperspectral data, and/or phenomic-data or combinations thereof” (Paragraph 48) and “In one embodiment, in the training stage, the multi-modal autoencoder, which includes an encoder trained to encode the input data (the features from the input data) obtained from the training and testing populations into a latent representation in the latent space and a decoder to decode the latent representation and to reconstruct the input data using unsupervised learning. In some aspects, the decoder is trained on existing genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data, or combinations thereof. The training of the autoencoder to learn the encoding and decoding of the two or more types of data or multi-omics data may be iterative until the likelihood of reconstructing the input data reaches a certain level or threshold of accuracy” (Paragraph 54). The examiner further notes that an autoencoder is trained via features of multiple modals of source multi-modal data. The examiner further notes that Jia teaches “generate a source task based on the features of individual single-modal signals” as “A VAE or other autoencoder, such as a MMAE or MVAE, may be trained with inputs from two or more types of multi-omics or multi-modal data. The inputs may include combinations of genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data. In some examples, the input data is genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, and/or phenomic, which includes but is not limited to genomic-wide data, exomic-wide data, epigenomic-wide data, transcriptomic-wide data, proteomic-wide data, metabolomic-wide data, hyperspectral data, and/or phenomic-data or combinations thereof” (Paragraph 48), “In one embodiment, in the training stage, the multi-modal autoencoder, which includes an encoder trained to encode the input data (the features from the input data) obtained from the training and testing populations into a latent representation in the latent space and a decoder to decode the latent representation and to reconstruct the input data using unsupervised learning. In some aspects, the decoder is trained on existing genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data, or combinations thereof. The training of the autoencoder to learn the encoding and decoding of the two or more types of data or multi-omics data may be iterative until the likelihood of reconstructing the input data reaches a certain level or threshold of accuracy” (Paragraph 54), and “a neural network or an autoencoder with an encoder and decoder is trained using unsupervised learning or any other suitable technique. Unsupervised learning is a method that may be used for training the autoencoder in a state in which a label is not allocated to training data. Unsupervised learning includes the task of trying to find hidden structure in unlabeled data. Some examples of unsupervised learning processes include but are not limited to: clustering (e.g., k-means, mixture models, hierarchical clustering), blind signal separation using feature extraction techniques for dimensionality reduction (e.g., principal component analysis, independent component analysis, non-negative matrix factorization, singular value decomposition) and artificial neural networks (e.g., self-organizing map, adaptive resonance theory). Clustering analysis is the assignment of a set of observations into subsets (called clusters) so that observations within the same cluster are similar according to some pre-designated criterion or criteria, while observations drawn from different clusters are dissimilar. Different clustering techniques make different assumptions on the structure of the data, often defined by some similarity metric and evaluated for example by internal compactness (similarity between members of the same cluster) and separation between different clusters. An unsupervised algorithm, e.g., a clustering or dimensionality reduction algorithm, may find previously unknown patterns in data sets without pre-existing labels. Accordingly, in at least one embodiment, the untrained autoencoder trains itself using unlabeled data. In some aspects, the unsupervised learning training dataset includes two or more types of input data, such as multi-omics or multi-modal data, without any associated output data. The training data set may be obtained from a training population, a testing population, or both. The autoencoder through training using unsupervised learning becomes capable of finding hidden structure in unlabeled data, learning groupings within training dataset, determining how individual inputs are related to the untrained dataset, or combinations thereof” (Paragraph 55). The examiner further notes that the iterative training of an autoencoder via unsupervised learning of features of multi-modal data entails the generation of the undefined source task in the broadest reasonable interpretation. Moreover, the instant specification simply defines a “task” as “an issue intended to address through machine learning or a task intended to perform through machine-learning” (Paragraph 42). The examiner further notes that Jia teaches “derive a learning method from the source task using the encoder” as “A VAE or other autoencoder, such as a MMAE or MVAE, may be trained with inputs from two or more types of multi-omics or multi-modal data. The inputs may include combinations of genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data. In some examples, the input data is genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, and/or phenomic, which includes but is not limited to genomic-wide data, exomic-wide data, epigenomic-wide data, transcriptomic-wide data, proteomic-wide data, metabolomic-wide data, hyperspectral data, and/or phenomic-data or combinations thereof” (Paragraph 48), “In one embodiment, in the training stage, the multi-modal autoencoder, which includes an encoder trained to encode the input data (the features from the input data) obtained from the training and testing populations into a latent representation in the latent space and a decoder to decode the latent representation and to reconstruct the input data using unsupervised learning. In some aspects, the decoder is trained on existing genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data, or combinations thereof. The training of the autoencoder to learn the encoding and decoding of the two or more types of data or multi-omics data may be iterative until the likelihood of reconstructing the input data reaches a certain level or threshold of accuracy” (Paragraph 54), and “a neural network or an autoencoder with an encoder and decoder is trained using unsupervised learning or any other suitable technique. Unsupervised learning is a method that may be used for training the autoencoder in a state in which a label is not allocated to training data. Unsupervised learning includes the task of trying to find hidden structure in unlabeled data. Some examples of unsupervised learning processes include but are not limited to: clustering (e.g., k-means, mixture models, hierarchical clustering), blind signal separation using feature extraction techniques for dimensionality reduction (e.g., principal component analysis, independent component analysis, non-negative matrix factorization, singular value decomposition) and artificial neural networks (e.g., self-organizing map, adaptive resonance theory). Clustering analysis is the assignment of a set of observations into subsets (called clusters) so that observations within the same cluster are similar according to some pre-designated criterion or criteria, while observations drawn from different clusters are dissimilar. Different clustering techniques make different assumptions on the structure of the data, often defined by some similarity metric and evaluated for example by internal compactness (similarity between members of the same cluster) and separation between different clusters. An unsupervised algorithm, e.g., a clustering or dimensionality reduction algorithm, may find previously unknown patterns in data sets without pre-existing labels. Accordingly, in at least one embodiment, the untrained autoencoder trains itself using unlabeled data. In some aspects, the unsupervised learning training dataset includes two or more types of input data, such as multi-omics or multi-modal data, without any associated output data. The training data set may be obtained from a training population, a testing population, or both. The autoencoder through training using unsupervised learning becomes capable of finding hidden structure in unlabeled data, learning groupings within training dataset, determining how individual inputs are related to the untrained dataset, or combinations thereof” (Paragraph 55). The examiner further notes that the iterative training of an autoencoder via unsupervised learning of features of multi-modal data entails the use of a derived “learning method” (which is undefined in the claims) for training/learning the autoencoder in and of itself. The examiner further notes that Jia teaches “perform the target task by training a model based on the features extracted by the encoder and the learning method” as “A VAE or other autoencoder, such as a MMAE or MVAE, may be trained with inputs from two or more types of multi-omics or multi-modal data. The inputs may include combinations of genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data. In some examples, the input data is genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, and/or phenomic, which includes but is not limited to genomic-wide data, exomic-wide data, epigenomic-wide data, transcriptomic-wide data, proteomic-wide data, metabolomic-wide data, hyperspectral data, and/or phenomic-data or combinations thereof” (Paragraph 48), “In one embodiment, in the training stage, the multi-modal autoencoder, which includes an encoder trained to encode the input data (the features from the input data) obtained from the training and testing populations into a latent representation in the latent space and a decoder to decode the latent representation and to reconstruct the input data using unsupervised learning. In some aspects, the decoder is trained on existing genomic, exomic, epigenomic, transcriptomic, proteomic, metabolomic, hyperspectral, or phenomic data, or combinations thereof. The training of the autoencoder to learn the encoding and decoding of the two or more types of data or multi-omics data may be iterative until the likelihood of reconstructing the input data reaches a certain level or threshold of accuracy” (Paragraph 54), and “a neural network or an autoencoder with an encoder and decoder is trained using unsupervised learning or any other suitable technique. Unsupervised learning is a method that may be used for training the autoencoder in a state in which a label is not allocated to training data. Unsupervised learning includes the task of trying to find hidden structure in unlabeled data. Some examples of unsupervised learning processes include but are not limited to: clustering (e.g., k-means, mixture models, hierarchical clustering), blind signal separation using feature extraction techniques for dimensionality reduction (e.g., principal component analysis, independent component analysis, non-negative matrix factorization, singular value decomposition) and artificial neural networks (e.g., self-organizing map, adaptive resonance theory). Clustering analysis is the assignment of a set of observations into subsets (called clusters) so that observations within the same cluster are similar according to some pre-designated criterion or criteria, while observations drawn from different clusters are dissimilar. Different clustering techniques make different assumptions on the structure of the data, often defined by some similarity metric and evaluated for example by internal compactness (similarity between members of the same cluster) and separation between different clusters. An unsupervised algorithm, e.g., a clustering or dimensionality reduction algorithm, may find previously unknown patterns in data sets without pre-existing labels. Accordingly, in at least one embodiment, the untrained autoencoder trains itself using unlabeled data. In some aspects, the unsupervised learning training dataset includes two or more types of input data, such as multi-omics or multi-modal data, without any associated output data. The training data set may be obtained from a training population, a testing population, or both. The autoencoder through training using unsupervised learning becomes capable of finding hidden structure in unlabeled data, learning groupings within training dataset, determining how individual inputs are related to the untrained dataset, or combinations thereof” (Paragraph 55). The examiner further notes that the iterative training of an autoencoder via unsupervised learning of features of multi-modal data entails the performing of the undefined target task in the broadest reasonable interpretation. Moreover, the instant specification simply defines a “task” as “an issue intended to address through machine learning or a task intended to perform through machine-learning” (Paragraph 42).
Jia does not explicitly teach:
G) wherein, to derive the learning method, the processor is configured to input the individual single-modal signals to a neural network to estimate a task intended to be performed; and
H) modulate the encoder based on the estimated task.
Kim, however, teaches “wherein, to derive the learning method, the processor is configured to input the individual single-modal signals to a neural network to estimate a task intended to be performed” as “The domain and task estimator 115 estimates domains and tasks of all pieces of the input support data D.sup.k′,t based on the embedded feature information according to the embedding result. In one embodiment, the domain and task estimator 115 may set the embedded feature information as an input of a multi-layer perceptron model and acquire the estimated domain and task of the support data as the output corresponding to the input. In this case, a dimension of an output stage for an output of the multi-layer perceptron model may be set to be smaller than that of an input stage for input” (Paragraphs 65-66) and “modulate the encoder based on the estimated task” as “The modulation information acquirer 120 acquires the modulation information of the initial parameter {tilde over (θ)}.sup.k of the task execution model based on the estimated domain and task. In one embodiment, the modulation information acquirer 120 may acquire the modulation information of the initial parameter {tilde over (θ)}.sup.k of the task execution model from the knowledge memory 130 using the estimated domain and task directly from the estimated domain and task or through a knowledge controller 125” (Paragraphs 67-68) and “The modulator 135 modulates the initial parameter {tilde over (θ)}.sup.k of the task execution model based on the modulation information. In this case, the modulator 135 may sum the modulation information directly acquired by the modulation information acquirer 120 and the modulation information acquired from the knowledge memory 130 by the knowledge controller 125 and may modulate the initial parameter {tilde over (θ)}.sup.k of the task execution model based on the summed modulation information” (Paragraph 73).
The examiner further notes that Kim teaches the estimation of a task to be performed via the input of support data to a model. The combination would result in inputting the single-model signals of Jia to estimate a task. Moreover, Kim teaches the modulation of parameters based on the estimated task. The combination would result in the modulation of the autoencoder of Jia via such modulated parameters.
It would have been obvious to one of ordinary skill in the art before the effective filing date of instant invention to combine the teachings of the cited references because teaching Kim’s would have allowed Jia’s to provide a method for coping with diverse domains and tasks, as noted by Kim (Paragraph 57).
10. Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Jia et al. (U.S. PGPUB 2023/0139567), in view of Kim et al. (U.S. PGPUB 2021/0398004) as applied to claims 1 and 11 above, and further in view of Li et al. (U.S. PGPUB 2023/0252295).
11. Regarding claim 8, Jia and Kim do not explicitly teach a method comprising:
A) wherein the multimodal dataset includes image, audio, and text signals.
Li, however, teaches “wherein the multimodal dataset includes image, audio, and text signals” as “an environment for providing environmental samples may include an environment in which various modalities blend with each other. According to different feature dimensions, various modalities may include but not be limited to at least one selected from: an audio modality, an image modality, a video modality, a text modality, or the like” (Paragraph 46) and “Through the above-mentioned embodiments of the present disclosure, the environmental samples used to train the model may be acquired based on a plurality of modalities” (Paragraph 51).
The examiner further notes that although Jia teaches the concept of training via multi-modal data, there is no explicit teaching that such data includes all of image, audio, and text signals. Nevertheless, Li teaches the concept of training a model via multimodal data constituting multiple different types of data (including image, audio, and text data). The combination would result in expanding Jia to have its multimodal data to include all of image, audio, and text data.
It would have been obvious to one of ordinary skill in the art before the effective filing date of instant invention to combine the teachings of the cited references because teaching Li’s would have allowed Jia’s and Kim’s to provide a method for achieving a good effect on model training, as noted by Li (Paragraph 51).
Regarding claim 18, Jia and Kim do not explicitly teach an apparatus comprising:
A) wherein the multimodal dataset includes image, audio, and text signals.
Li, however, teaches “wherein the multimodal dataset includes image, audio, and text signals” as “an environment for providing environmental samples may include an environment in which various modalities blend with each other. According to different feature dimensions, various modalities may include but not be limited to at least one selected from: an audio modality, an image modality, a video modality, a text modality, or the like” (Paragraph 46) and “Through the above-mentioned embodiments of the present disclosure, the environmental samples used to train the model may be acquired based on a plurality of modalities” (Paragraph 51).
The examiner further notes that although Jia teaches the concept of training via multi-modal data, there is no explicit teaching that such data includes all of image, audio, and text signals. Nevertheless, Li teaches the concept of training a model via multimodal data constituting multiple different types of data (including image, audio, and text data). The combination would result in expanding Jia to have its multimodal data to include all of image, audio, and text data.
It would have been obvious to one of ordinary skill in the art before the effective filing date of instant invention to combine the teachings of the cited references because teaching Li’s would have allowed Jia’s and Kim’s to provide a method for achieving a good effect on model training, as noted by Li (Paragraph 51).
Allowable Subject Matter
12. Claim 2 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Dependent claims 3 and 9-10 are deemed allowable for depending on the deemed allowable subject matter of dependent claim 2.
Claim 4 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim 5 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Dependent claims 6-7 are deemed allowable for depending on the deemed allowable subject matter of dependent claim 5.
Claim 12 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Dependent claims 13 and 19-20 are deemed allowable for depending on the deemed allowable subject matter of dependent claim 12.
Claim 14 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim 15 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Dependent claims 16-17 are deemed allowable for depending on the deemed allowable subject matter of dependent claim 15.
Response to Arguments
13. Applicant’s arguments with respect to claims 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument (See newly applied art of Kim).
Conclusion
14. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
U.S. PGPUB 2019/0197366 issued to Kecskemethy et al. on 27 June 2019. The subject matter disclosed therein is pertinent to that of claims 1-20 (e.g., methods to process multi-modal data).
15. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Contact Information
16. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Mahesh Dwivedi whose telephone number is (571) 272-2731. The examiner can normally be reached on Monday to Friday 8:20 am – 4:40 pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Charles Rones can be reached (571) 272-4085. The fax number for the organization where this application or proceeding is assigned is (571) 273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
Mahesh Dwivedi
Primary Examiner
Art Unit 2168
September 20, 2026
/MAHESH H DWIVEDI/Primary Examiner, Art Unit 2168