DETAILED ACTION
This action is in response to the amendment filed 07/29/2026. Claims 1, 4-5, 7, 9, 12, 15, 23, 28, 30-31, and 38-46 are pending and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1, 4-5, 7, 9, 12, 15, 23, 28, 30-31, and 38-46 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites the limitation "in response to all features of the original training dataset having been reproduced with synthetic data in the synthetic training dataset" in lines 17-18. There is insufficient antecedent basis for this limitation in the claim. Specifically, claim 1 does not recite reproducing the original dataset with the synthetic training dataset. Instead, claim 1 recites “inserting a synthetic feature c’I corresponding to the estimate y’I of the target vector yi into the synthetic training dataset.” For purposes of examination, Examiner has interpreted this reproduction to be generating a synthetic training dataset as recited in the preamble of the claim.
Regarding claims 4-5, 7, 9, 12, 15, and 43-46, claims 4-5, 7, 9, 12, 15, and 43-46 are rejected for at least the same reasons as claim 1 since claims 4-5, 7, 9, 12, 15, and 43-46 depend on claim 1.
Claim 23 recites the limitation "in response to all features of the original training dataset having been reproduced with synthetic data in the synthetic training dataset" in lines 18-19. There is insufficient antecedent basis for this limitation in the claim. Specifically, claim 23 does not recite reproducing the original dataset with the synthetic training dataset. Instead, claim 23 recites “inserting a synthetic feature c’I corresponding to the estimate y’I of the target vector yi into the synthetic training dataset.” For purposes of examination, Examiner has interpreted this reproduction to be generating a synthetic training dataset.
Regarding claims 38-42, claims 38-42 are rejected for at least the same reasons as claim 23 since claims 38-42 depend on claim 23.
Claim 28 recites the limitation "in response to all features of the original training dataset having been reproduced with synthetic data in the synthetic training dataset" in lines 19-20. There is insufficient antecedent basis for this limitation in the claim. Specifically, claim 28 does not recite reproducing the original dataset with the synthetic training dataset. Instead, claim 28 recites “inserting a synthetic feature c’I corresponding to the estimate y’I of the target vector yi into the synthetic training dataset.” For purposes of examination, Examiner has interpreted this reproduction to be generating a synthetic training dataset as recited in the preamble of the claim.
Regarding claims 30-31, claims 30-31 are rejected for at least the same reasons as claim 28 since claims 30-31 depend on claim 28.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 1, 4-5, 7, 9, 23, 38-44 and 46 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wan et al. (“Improving protein function prediction with synthetic feature samples created by generative adversarial networks”) (hereafter referred to as Wan) in view of Cimentada (“The LOO and the Bootstrap”) (hereafter referred to as Cimentada) in further view of Szeto et al. (US 2018/0018590 A1) (hereafter referred to as Szeto) and Hardy et al. (“MD-GAN: Multi-Discriminator Generative Adversarial Networks for Distributed Datasets”) (hereafter referred to as Hardy).
Regarding claim 1, Wan teaches
A method of generating a synthetic training dataset for training a machine learning model using an original training dataset including a plurality of features (Wan, page 1, abstract, “In this work, we propose a novel generative adversarial networks-based method, namely FFPred-GAN, to accurately learn the high-dimensional distributions of protein sequence-based biophysical features and also generate high-quality synthetic protein feature samples.” Examiner notes that the protein sequence-based biophysical features are the original dataset and the synthetic training dataset is the synthetic protein feature samples.
In response to all features of the original training dataset having been reproduced with synthetic data in the synthetic training dataset, appending the synthetic training dataset to the original training dataset to form a hybrid training dataset having a larger size than the original training dataset; and training the machine learning model using the hybrid training dataset to increase an initial training accuracy of the machine learning model relative to training the machine learning model only using the original training dataset (Wan, page 3, 2nd paragraph, “On the last step, FFPred-GAN uses the Classifier Two-Sample Tests (CTST) to select the optimal synthetic training protein feature samples, which are used to augment the original training samples. During the down-stream machine learning classifier training stage, the optimal synthetic samples are expected to derive better classifiers, leading to higher predictive accuracy” where “the SVM trained by the augmented training protein feature samples learning those decision boundaries that successfully separate the protein samples distributed on the right corner of the figure” (Wan, page 9, last paragraph) and Wan, page 3, Figure 1
PNG
media_image1.png
338
748
media_image1.png
Greyscale
Examiner notes that the synthetic training dataset is the optimal synthetic training protein feature samples, the training dataset is the original training samples, and the hybrid training dataset is the augmented training protein feature samples. Examiner further notes that the augmenting the datasets is appending the synthetic training dataset to the training dataset to form a hybrid training dataset. Examiner also notes that the SVM is trained by the augmented training protein feature samples or the hybrid training dataset.) Examiner further notes that the hybrid dataset is shown in Figure 1 by combining the real and synthetic features into the critic which is larger than just the real feature samples alone. Examiner additionally notes that the higher predictive accuracy is an increase to an initial training accuracy.)
Wan does not teach, but Cimentada does teach
selecting a feature ci of the original training dataset as a target vector yi (Cimentada, page 1, 2nd paragraph, “Let’s imagine a data set with 30 rows. We separate the 1st row to be the test data.” Examiner notes that the first row is the target vector.);
selecting remaining features of the original training dataset as a set of training input vectors X\i, where X\i includes all features of the original training dataset other than a feature corresponding to the selected feature ci (Cimentada, page 1, 2nd paragraph, “We separate the 1st row to be the test data and the remaining 29 rows to be the training data.” Examiner notes that the training input vectors are the remaining 29 rows.);
training a prediction model f(yi|X\i) (Cimentada, page 1, 2nd paragraph, “We fit the model on the training data and then predict the one observation we left out.” Examiner notes that fitting the model is training a prediction model.);
generating an estimate y'i of the target vector yi by applying the prediction model to the set of training vectors X\i (Cimentada, page 1, 2nd paragraph, “We fit the model on the training data and then predict the one observation we left out.” Examiner notes that predicting the one observation left out is generating an estimate of the target vector.);
repeating, for a plurality of features of the original training dataset, operations of selecting a feature of the original training dataset, selecting remaining features of the original training dataset, training the prediction model, generating the estimate of the target vector (Cimentada, page 1, 2nd paragraph, “LOOCV: Let’s imagine a data set with 30 rows. We separate the 1st row to be the test data and the remaining 29 rows to be the training data. We fit the model on the training data and then predict the one observation we left out. We record the model accuracy and then repeat but predicting the 2nd row from training the model on row 1 and 3:30. We repeat until every row has been predicted.” Examiner notes that repeating until every row has been predicted is repeating for a plurality of features of the original training dataset, operations of selecting a feature of the original training dataset, selecting remaining features of the original training dataset, training the prediction mode, and generating the estimate of the target vector.)
Wan and Cimentada are considered analogous to the claimed invention because they both use the Leave One Out method on generated data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan to select features, train a model, and generate an estimate like in Cimentada. Doing so would be advantageous because “it uses all the data. At some point, every rows gets to be the test set and training set, maximizing information. In fact, it uses almost ALL the data as the original data set as the training set is just N-1” (Cimentada, page 2, bullet points).
Wan in view of Cimentada teach the estimate y'i of the target vector yi, but does not teach inserting a synthetic feature corresponding to the estimate into a synthetic training dataset. Szeto does teach
and inserting a synthetic feature c'i corresponding to the estimate y'i …into the synthetic training dataset (Szeto, page 9, paragraph 0018, “The modeling engine further generates one or more private data distributions from the local private data training set where the private data distributions represent the nature of the local private data used to create the trained model. The modeling engine uses the private data distributions to generate a set of proxy data, which can be considered synthetic data or Monte Carlo data having the same general data distribution characteristics as the local private data, while also lacking the actual private or restricted features of the local, private data….The modeling engine then attempts to validate that the set of proxy data is a reasonable training set stand-in for the local, private data by creating a trained proxy model from the set of proxy data” where “From the private data distributions, the machine learning engine can identify or otherwise calculate one or more salient private data features that describe the nature of the private data distributions” (Szeto, page 10, paragraph 0019) and where “private data servers transmit salient features of aggregated private data to a non-private computing devices which in turn creates proxy data for integrating into a trained global model” (Szeto, page 10, paragraph 0026). Examiner notes that the proxy data is the synthetic training dataset. Examiner further notes that the estimate is the salient private data features. By creating proxy data from the salient private data features, the synthetic features correspond to the estimate.)
repeating … inserting the synthetic feature into the synthetic training dataset (Szeto, page 24, paragraph 0118, “then the global modeling engine can repeat operations 660 through 680 until a satisfactory similar trained proxy model is generated” where “operation 660 shifts focus from the modeling engine in an entity’s private data server to the non-private computing device’s global modeling engine (see FIG. 1, global modeling engine 136). The global modeling engine receives the salient private data features and locally re-instantiates the private data features distributions in memory. As discussed previously with respect to operation 540, the global modeling engine generates proxy data from the salient private data features, for example, by using the re-instantiated private data distributions as probability distributions to generate new, synthetic sample data” (Szeto, page 23, paragraph 01115). Examiner notes that the repeating operation 660 is repeating inserting the synthetic feature into the synthetic training dataset.)
Wan in view of Cimentada and Szeto are analogous to the claimed invention because they generate synthetic data to train machine learning models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan in view of Cimentada to insert synthetic data into a dataset like in Szeto. Doing so “is considered advantageous because it provides for generating synthetic data capable of reproducing the knowledge gained from private data” (Szeto, page 18, paragraph 0078).
Wan in view of Cimentada and Szeto do not explicitly disclose, but Hardy does disclose
and transmitting weights of the machine learning model, trained using the hybrid training dataset, to a master node (Hardy, page 3, 2nd column, 1st paragraph, “Workers perform iterations locally on their data and every E epochs (i.e., each worker passes E times the data in their GAN) they send the resulting parameters to the server” where “The server generates a set K of k batches K = {X(1),…, X(k)}, with k ≤ N. Each X(i) is composed of b data generated by G. The server then selects, for each worker n, two distinct batches, say X(i) and X(j), which are sent to worker n and locally renamed as
X
n
(
g
)
and
X
n
(
d
)
. The way in which the two distinct batches are selected is discussed in Section IV-B1. Each worker n performs L learning iterations on its discriminator Dn (see Section II-1) using
X
n
(
d
)
and
X
n
(
r
)
, where
X
n
(
r
)
is a batch of real data extracted locally from Bn” (Hardy, page 3, 2nd column, first two bullet points) Examiner notes that performing iterations using both real and generated data is appending the synthetic training dataset to the training dataset to form a hybrid dataset and training the model. Examiner further notes that the indication from the master node is the server sending the generated data to the worker. Examiner additionally notes that the server is the master node and the workers transmit parameters or trained weights.)
Wan in view of Cimentada, Szeto and Hardy are analogous to the claimed invention because they generate synthetic data to train machine learning models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan in view of Cimentada and Szeto to use a federated learning system to generate synthetic data to train a model like in Hardy. Doing so “address[es] the problem of distributing GANs so that they are able to train over datasets that are spread on multiple workers” (Hardy, page 1, abstract).
Regarding claim 4, Wan in view of Cimentada, Szeto and Hardy teach the method of Claim 1. Wan in view of Cimentada and Szeto further teach
wherein the prediction model comprises a bagging or boosting algorithm (Szeto, page 16, paragraph 0069, “More specifically, machine learning algorithms 295 can include implementations of one or more of the following algorithms, … a boosting algorithm…bootstrapped aggregation (bagging).”).
Wan in view of Cimentada and Szeto are analogous to the claimed invention because they generate synthetic data to train prediction models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have implemented the prediction model in Wan in view of Cimentada to use a bagging or boosting algorithm like in Szeto. Thus, this would be applying a known technique (bagging or boosting algorithm) to a known device (prediction model) ready for improvement to yield predictable results (feature prediction) (MPEP 2143 I. (C) Use of known technique to improve similar devices (methods, or products) in the same way).
Regarding claim 5, Wan in view of Cimentada, Szeto and Hardy teach the method of Claim 1. Wan in view of Cimentada and Szeto further teach
wherein the prediction model comprises a random forest prediction model or gradient boosting tree model (Szeto, page 16, paragraph 0069, “More specifically, machine learning algorithms 295 can include implementations of one or more of the following algorithms, … gradient boosted regression trees (GBRT), a random forest.”).
Wan in view of Cimentada and Szeto are analogous to the claimed invention because they generate synthetic data to train prediction models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have implemented the prediction model in Wan in view of Cimentada to be a random forest prediction model or gradient boosting tree model like in Szeto. Thus, this would be applying a known technique (bagging or boosting algorithm) to a known device (prediction model) ready for improvement to yield predictable results (feature prediction) (MPEP 2143 I. (C) Use of known technique to improve similar devices (methods, or products) in the same way).
Regarding claim 7, Wan in view of Cimentada, Szeto and Hardy teach the method of Claim 1. Wan in view of Cimentada further teach
wherein generating the estimate y'i of the target vector yi comprises: generating an estimate y'i of the target vector yi by applying the prediction model as f(X\i) - > y'i (Cimentada, page 1, 2nd paragraph, “We fit the model on the training data and then predict the one observation we left out.” Examiner notes that predicting the one observation left out is generating an estimate of the target vector. Examiner further notes that predicting after fitting the model on the training data is applying the prediction model where the training vectors map to the estimate.).
Wan and Cimentada are considered analogous to the claimed invention because they both use the Leave One Out method on generated data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan to repeat the steps of select features, train a model, and generate an estimate like in Cimentada. Doing so would be advantageous because “it uses all the data. At some point, every rows gets to be the test set and training set, maximizing information. In fact, it uses almost ALL the data as the original data set as the training set is just N-1” (Cimentada, page 2, bullet points).
Regarding claim 9, Wan in view of Cimentada, Szeto and Hardy teach the method of Claim 1. Wan further teaches
wherein the machine learning model comprises a neural network (Wan, page 13, 1st paragraph “Wasserstein generative adversarial networks with gradient penalty are a type of Generative Adversarial Networks (GANs), which are well-known to be highly capable of learning high-dimensional distributions from data samples. In general, conventional GANs are composed of two neural networks, i.e. the generator G and the discriminator (a.k.a. critic) D” where “in this work, we use the generator of well-trained WGAN-GP [Wasserstein generative adversarial networks with gradient penalty] models to generate synthetic samples” (Wan, page 13, last paragraph).)
Regarding claim 23, Wan teaches
A computing device, comprising processing circuitry (Wan, page 11, 1st paragraph, We further discuss the computational time cost (i.e. the actual running time obtained by using CPU-based PyTorch with a standard Linux computing cluster) and the training sample sizes (i.e. the number of training protein feature samples) for running FFPred-GAN to generate the optimal synthetic protein feature samples for individual GO terms.” Examiner notes that the CPU is the computing device and processing circuitry.)
In response to all features of the original training dataset having been reproduced with synthetic data in the synthetic training dataset, appending the synthetic training dataset to the original training dataset to form a hybrid training dataset having a larger size than the original training dataset; and training the machine learning model using the hybrid training dataset to increase an initial training accuracy of the machine learning model relative to training the machine learning model only using the original training dataset (Wan, page 3, 2nd paragraph, “On the last step, FFPred-GAN uses the Classifier Two-Sample Tests (CTST) to select the optimal synthetic training protein feature samples, which are used to augment the original training samples. During the down-stream machine learning classifier training stage, the optimal synthetic samples are expected to derive better classifiers, leading to higher predictive accuracy” where “the SVM trained by the augmented training protein feature samples learning those decision boundaries that successfully separate the protein samples distributed on the right corner of the figure” (Wan, page 9, last paragraph) and Wan, page 3, Figure 1
PNG
media_image1.png
338
748
media_image1.png
Greyscale
Examiner notes that the synthetic training dataset is the optimal synthetic training protein feature samples, the training dataset is the original training samples, and the hybrid training dataset is the augmented training protein feature samples. Examiner further notes that the augmenting the datasets is appending the synthetic training dataset to the training dataset to form a hybrid training dataset. Examiner also notes that the SVM is trained by the augmented training protein feature samples or the hybrid training dataset.) Examiner further notes that the hybrid dataset is shown in Figure 1 by combining the real and synthetic features into the critic which is larger than just the real feature samples alone. Examiner additionally notes that the higher predictive accuracy is an increase to an initial training accuracy.)
Wan does not teach, but Cimentada does teach
selecting a feature ci of a training dataset as a target vector yi , the training dataset comprising a plurality of features (Cimentada, page 1, 2nd paragraph, “Let’s imagine a data set with 30 rows. We separate the 1st row to be the test data.” Examiner notes that the first row is the target vector and each row is a feature.);
selecting remaining features of the training dataset as a set of training vectors X\i, where X\i includes all features of the training dataset other than feature ci (Cimentada, page 1, 2nd paragraph, “We separate the 1st row to be the test data and the remaining 29 rows to be the training data.” Examiner notes that the training input vectors are the remaining 29 rows.);
training a prediction model f(yi|X\i) (Cimentada, page 1, 2nd paragraph, “We fit the model on the training data and then predict the one observation we left out.” Examiner notes that fitting the model is training a prediction model.);
generating an estimate y'i of the target vector yi by applying the prediction model to the set of training vectors X\i (Cimentada, page 1, 2nd paragraph, “We fit the model on the training data and then predict the one observation we left out.” Examiner notes that predicting the one observation left out is generating an estimate of the target vector.);
repeating, for a plurality of features of the original training dataset, operations of selecting a feature of the original training dataset, selecting remaining features of the original training dataset, training the prediction model, generating the estimate of the target vector (Cimentada, page 1, 2nd paragraph, “LOOCV: Let’s imagine a data set with 30 rows. We separate the 1st row to be the test data and the remaining 29 rows to be the training data. We fit the model on the training data and then predict the one observation we left out. We record the model accuracy and then repeat but predicting the 2nd row from training the model on row 1 and 3:30. We repeat until every row has been predicted.” Examiner notes that repeating until every row has been predicted is repeating for a plurality of features of the original training dataset, operations of selecting a feature of the original training dataset, selecting remaining features of the original training dataset, training the prediction mode, and generating the estimate of the target vector.)
Wan and Cimentada are considered analogous to the claimed invention because they both use the Leave One Out method on generated data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan to select features, train a model, and generate an estimate like in Cimentada. Doing so would be advantageous because “it uses all the data. At some point, every rows gets to be the test set and training set, maximizing information. In fact, it uses almost ALL the data as the original data set as the training set is just N-1” (Cimentada, page 2, bullet points).
Wan in view of Cimentada teach the estimate y'i of the target vector yi, but does not teach inserting a synthetic feature corresponding to the estimate into a synthetic training dataset. Szeto does teach
memory coupled to the processing circuitry and including instructions that are executable by the processing circuitry to cause the computer device to perform operations comprising (Szeto, page 9, paragraph 0018, “The private data servers are computing devices having one or more processors that are configurable to execute software instructions stored in a non-transitory computer readable memory, where execution of the software instructions gives rise to a modeling engine on the private data server.” ):
and inserting a synthetic feature c'i corresponding to the estimate y'i of the target vector yi into a synthetic training dataset (Szeto, page 9, paragraph 0018, “The modeling engine further generates one or more private data distributions from the local private data training set where the private data distributions represent the nature of the local private data used to create the trained model. The modeling engine uses the private data distributions to generate a set of proxy data, which can be considered synthetic data or Monte Carlo data having the same general data distribution characteristics as the local private data, while also lacking the actual private or restricted features of the local, private data….The modeling engine then attempts to validate that the set of proxy data is a reasonable training set stand-in for the local, private data by creating a trained proxy model from the set of proxy data” where “From the private data distributions, the machine learning engine can identify or otherwise calculate one or more salient private data features that describe the nature of the private data distributions” (Szeto, page 10, paragraph 0019) and where “private data servers transmit salient features of aggregated private data to a non-private computing devices which in turn creates proxy data for integrating into a trained global model” (Szeto, page 10, paragraph 0026). Examiner notes that the proxy data is the synthetic training dataset. Examiner further notes that the estimate of the target vector is the salient private data features. By creating proxy data from the salient private data features, the synthetic features correspond to the estimate.).
repeating … inserting the synthetic feature into the synthetic training dataset (Szeto, page 24, paragraph 0118, “then the global modeling engine can repeat operations 660 through 680 until a satisfactory similar trained proxy model is generated” where “operation 660 shifts focus from the modeling engine in an entity’s private data server to the non-private computing device’s global modeling engine (see FIG. 1, global modeling engine 136). The global modeling engine receives the salient private data features and locally re-instantiates the private data features distributions in memory. As discussed previously with respect to operation 540, the global modeling engine generates proxy data from the salient private data features, for example, by using the re-instantiated private data distributions as probability distributions to generate new, synthetic sample data” (Szeto, page 23, paragraph 01115). Examiner notes that the repeating operation 660 is repeating inserting the synthetic feature into the synthetic training dataset.)
Wan in view of Cimentada and Szeto are analogous to the claimed invention because they generate synthetic data to train machine learning models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan in view of Cimentada to insert synthetic data into a dataset like in Szeto. Doing so “is considered advantageous because it provides for generating synthetic data capable of reproducing the knowledge gained from private data” (Szeto, page 18, paragraph 0078). Additionally, it would have been obvious to a person having ordinary skill in the art prior to the effective filing date to have implemented Wan in view of Cimentada on the computing device in Szeto. Thus, this would be applying a known technique (generating synthetic data) to a known device (memory coupled to the processing circuitry and including instructions) ready for improvement to yield predictable results (train machine learning models) (MPEP 2143 I. (C) Use of known technique to improve similar devices (methods, or products) in the same way).
Wan in view of Cimentada and Szeto do not explicitly disclose, but Hardy does disclose
and transmitting weights of the machine learning model, trained using the hybrid training dataset, to a master node (Hardy, page 3, 2nd column, 1st paragraph, “Workers perform iterations locally on their data and every E epochs (i.e., each worker passes E times the data in their GAN) they send the resulting parameters to the server” where “The server generates a set K of k batches K = {X(1),…, X(k)}, with k ≤ N. Each X(i) is composed of b data generated by G. The server then selects, for each worker n, two distinct batches, say X(i) and X(j), which are sent to worker n and locally renamed as
X
n
(
g
)
and
X
n
(
d
)
. The way in which the two distinct batches are selected is discussed in Section IV-B1. Each worker n performs L learning iterations on its discriminator Dn (see Section II-1) using
X
n
(
d
)
and
X
n
(
r
)
, where
X
n
(
r
)
is a batch of real data extracted locally from Bn” (Hardy, page 3, 2nd column, first two bullet points) Examiner notes that performing iterations using both real and generated data is appending the synthetic training dataset to the training dataset to form a hybrid dataset and training the model. Examiner further notes that the indication from the master node is the server sending the generated data to the worker. Examiner additionally notes that the server is the master node and the workers transmit parameters or trained weights.)
Wan in view of Cimentada, Szeto and Hardy are analogous to the claimed invention because they generate synthetic data to train machine learning models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan in view of Cimentada and Szeto to use a federated learning system to generate synthetic data to train a model like in Hardy. Doing so “address[es] the problem of distributing GANs so that they are able to train over datasets that are spread on multiple workers” (Hardy, page 1, abstract).
Regarding claim 38, claim 38 recites substantially similar limitations to claim 4, and is therefore rejected under the same analysis.
Regarding claim 39, claim 39 recites substantially similar limitations to claim 5, and is therefore rejected under the same analysis.
Regarding claim 40, claim 40 recites substantially similar limitations to claim 7, and is therefore rejected under the same analysis.
Regarding claim 41, claim 41 recites substantially similar limitations to claim 9, and is therefore rejected under the same analysis.
Regarding claim 42, Wan in view of Cimentada and Szeto teach the computing device of Claim 23. Wan in view of Cimentada and Szeto further teach appending the synthetic training dataset to the training dataset to form the hybrid training dataset and training the machine learning model as shown in claim 1. Wan in view of Cimentada and Szeto does not teach performing the appending the datasets and training the model in response to an indication from a master node in a federated learning system. However, Hardy does teach
wherein appending the synthetic training dataset to the original training dataset to form the hybrid training dataset and training the machine learning model are performed in response to an indication from a master node in a federated learning system (Hardy, page 3, 2nd column, first two bullet points, “The server generates a set K of k batches K = {X(1),…, X(k)}, with k ≤ N. Each X(i) is composed of b data generated by G. The server then selects, for each worker n, two distinct batches, say X(i) and X(j), which are sent to worker n and locally renamed as
X
n
(
g
)
and
X
n
(
d
)
. The way in which the two distinct batches are selected is discussed in Section IV-B1. Each worker n performs L learning iterations on its discriminator Dn (see Section II-1) using
X
n
(
d
)
and
X
n
(
r
)
, where
X
n
(
r
)
is a batch of real data extracted locally from Bn.” Examiner notes that performing iterations using both real and generated data is appending the synthetic training dataset to the training dataset to form a hybrid dataset and training the model. Examiner further notes that the indication from the master node is the server sending the generated data to the worker.).
Wan in view of Cimentada, Szeto and Hardy are analogous to the claimed invention because they generate synthetic data to train machine learning models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan in view of Cimentada and Szeto to use a federated learning system to generate synthetic data to train a model like in Hardy. Doing so “address[es] the problem of distributing GANs so that they are able to train over datasets that are spread on multiple workers” (Hardy, page 1, abstract).
Regarding claim 43, Wan in view of Cimentada, Szeto and Hardy teach the method of Claim 1. Wan in view of Cimentada, Szeto, and Hardy further teaches
training a set of weights for a local machine learning model using a real training dataset (Hardy, page 3, 2nd column, 1st paragraph, “Workers perform iterations locally on their data and every E epochs (i.e., each worker passes E times the data in their GAN) they send the resulting parameters to the server” where “The server generates a set K of k batches K = {X(1),…, X(k)}, with k ≤ N. Each X(i) is composed of b data generated by G. The server then selects, for each worker n, two distinct batches, say X(i) and X(j), which are sent to worker n and locally renamed as
X
n
(
g
)
and
X
n
(
d
)
. The way in which the two distinct batches are selected is discussed in Section IV-B1. Each worker n performs L learning iterations on its discriminator Dn (see Section II-1) using
X
n
(
d
)
and
X
n
(
r
)
, where
X
n
(
r
)
is a batch of real data extracted locally from Bn” (Hardy, page 3, 2nd column, first two bullet points) Examiner notes that the weights are the parameters, and the real training dataset is the batches of data.);
transmitting the set of weights of the local machine learning model to the master node (Hardy, page 3, 2nd column, 1st paragraph, “Workers perform iterations locally on their data and every E epochs (i.e., each worker passes E times the data in their GAN) they send the resulting parameters to the server” where “The server generates a set K of k batches K = {X(1),…, X(k)}, with k ≤ N. Each X(i) is composed of b data generated by G. The server then selects, for each worker n, two distinct batches, say X(i) and X(j), which are sent to worker n and locally renamed as
X
n
(
g
)
and
X
n
(
d
)
. The way in which the two distinct batches are selected is discussed in Section IV-B1. Each worker n performs L learning iterations on its discriminator Dn (see Section II-1) using
X
n
(
d
)
and
X
n
(
r
)
, where
X
n
(
r
)
is a batch of real data extracted locally from Bn” (Hardy, page 3, 2nd column, first two bullet points) Examiner notes that the server is the master node and the workers transmit parameters or trained weights.);
and receiving, from the master node, a message with instructions to generate synthetic training data (Hardy, page 4, Figure 1b (see image below)
PNG
media_image2.png
400
626
media_image2.png
Greyscale
Examiner notes that the server sends the generator to the workers which maps to receiving a message to generate synthetic training data.)
Wan in view of Cimentada, Szeto and Hardy are analogous to the claimed invention because they generate synthetic data to train machine learning models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan in view of Cimentada and Szeto to use a federated learning system to generate synthetic data to train a model like in Hardy. Doing so “address[es] the problem of distributing GANs so that they are able to train over datasets that are spread on multiple workers” (Hardy, page 1, abstract).
Regarding claim 44, Wan in view of Cimentada, Szeto and Hardy teach the method of Claim 1. Wan in view of Cimentada, Szeto, and Hardy further teach
generating a quality metric measuring a quality of the synthetic training dataset (Hardy, page 5, 2nd column, 3rd paragraph, “Each worker n hosts a discriminator Dn and a training dataset Bn. It receives batches of generated imaged split into two parts Xn(d) and Xn(g). The generated images Xn(d) are used for training Dn to discriminate those generated images from real images. The learning is performed as a classical deep learning operation on a standlone server [1]. A worker n computes the gradient
∆
θ
n
of the error function Jdisc applied to the batch of generated images Xn(d) and a batch or real image Xn(r) taken from Bn. As indicated in Section II-1, this operation is iterated L times. The second batch Xn(g) of generated images is used to compute the error term F-n of generator G. Once computed, Fn is sent to the server for computation of gradients
∆
w
.” Examiner notes that the quality metric is the error term.);
Transmitting the quality metric to the master node(Hardy, page 5, 2nd column, 3rd paragraph, “Each worker n hosts a discriminator Dn and a training dataset Bn. It receives batches of generated imaged split into two parts Xn(d) and Xn(g). The generated images Xn(d) are used for training Dn to discriminate those generated images from real images. The learning is performed as a classical deep learning operation on a standlone server [1]. A worker n computes the gradient
∆
θ
n
of the error function Jdisc applied to the batch of generated images Xn(d) and a batch or real image Xn(r) taken from Bn. As indicated in Section II-1, this operation is iterated L times. The second batch Xn(g) of generated images is used to compute the error term F-n of generator G. Once computed, Fn is sent to the server for computation of gradients
∆
w
.” Examiner notes that the quality metric is the error term and the master node is the server.);
And receiving, from the master node, instructions to proceed with training the machine learning model using the synthetic training dataset (Hardy, page 5, 2nd column, 3rd paragraph, “Each worker n hosts a discriminator Dn and a training dataset Bn. It receives batches of generated imaged split into two parts Xn(d) and Xn(g). The generated images Xn(d) are used for training Dn to discriminate those generated images from real images. The learning is performed as a classical deep learning operation on a standlone server [1]. A worker n computes the gradient
∆
θ
n
of the error function Jdisc applied to the batch of generated images Xn(d) and a batch or real image Xn(r) taken from Bn. As indicated in Section II-1, this operation is iterated L times. The second batch Xn(g) of generated images is used to compute the error term F-n of generator G. Once computed, Fn is sent to the server for computation of gradients
∆
w
.” Examiner notes that the quality metric is the error term and the master node is the server. Examiner further notes that by sending the error term to the server, the master node is receiving instructions to proceed with training since the error term is used further for the computation of gradients.);
Regarding claim 46, Wan in view of Cimentada, Szeto and Hardy teach the method of Claim 1. Wan further teaches
Wherein the synthetic training dataset appended to the training dataset creates the hybrid training dataset that doubles a size of the original training dataset (Wan, page 14, 2nd paragraph, “More specifically, given two equal-sized sets of samples respectively following two distributions P and Q, the CTST consider to accept or reject a null hypothesis of P being not equal to Q. If the null hypothesis is rejected, the classification accuracy on predicting the binary labels of held-out samples will be near the chance-level (i.e. 50.0%). Therefore, in terms of a metric evaluating the quality of generated synthetic samples, a classification accuracy of 100.0% means the synthetic samples are of poor quality, due to the fact that the synthetic samples are significantly different to the real ones. Analogously, a classification accuracy of 0.0% also suggests poor quality of the synthetic samples, due to the fact that the real and synthetic samples appear identical, suggesting that the model is merely regenerating the training samples and has failed to generate diverse new synthetic samples.” Examiner notes that P and Q are the real and synthetic training datasets that are combined in the critic to create a hybrid training dataset. Examiner further notes that since P and Q are equal, the size of the original training dataset is doubled when Q is added.)
Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wan in view of Cimentada in further view of Szeto, Hardy, and He et al. (US 2020/0386811 A1) (hereafter referred to as He).
Regarding claim 12, Wan in view of Cimentada, Szeto, and Hardy teach the method of Claim 1. Wan further teaches
providing a preliminary training dataset (Wan, page 3, 2nd paragraph, “To begin with, FFPred-GAN adopts the widely-used FFPred feature extractor to derive protein biophysical information based on the raw amino acid sequences.” Examiner notes that the raw amino acid sequences are the preliminary dataset.);
splitting the preliminary training dataset into the original training dataset and a verification dataset before generating the synthetic training dataset (Wan, page 14, last paragraph, “The protein set for each GO term was further split into the training and testing protein sets with a proportion of 7:3. The total 258 dimensions of protein sequence-derived biophysical features…are used to describe the proteins.” Examiner notes that the protein set is the preliminary training dataset.);
performing feature reduction on the preliminary training dataset before splitting the preliminary training dataset into the original training dataset and the verification dataset (Wan, page 3, 2nd paragraph, “To begin with, FFPred-GAN adopts the widely-used FFPred feature extractor to derive protein biophysical information based on the raw amino acid sequences. For each protein sequence, 258 dimensional features are generated to describe 13 groups of protein biophysical information, such as secondary structure, amino acid composition and presence of motifs” where “we train two FFPred-GAN models for each GO term by using two different sets of protein samples with different class labels” (Wan, page 3, last paragraph) and “the protein set for each GO term was further split into the training and testing protein sets with a proportion of 7:3” (Wan, page 14, last paragraph). Examiner notes that the raw amino acid sequences are the preliminary dataset and the feature extractor is feature reduction. Examiner further notes that the feature extraction occurs and produces protein samples or protein sets. After the protein samples are extracted, the samples are split into training and verification datasets.);
Wan in view of Cimentada does not teach, but Szeto does teach
verifying the machine learning model using the verification dataset (Szeto, page 21, paragraph 0098, “In other embodiments, proxy data 422 can be partitioned into training and validation data sets, which can then be used for cross-fold validation….the validation proxy data would be provided to the trained proxy model for validation” and “The model instructions can be considered as one or more command that instruct the modeling engine to use at least some of the local private data in order to create a trained actual model according to an implementation of a machine learning algorithm (e.g., support vector machine, neural network)” (Szeto, page 9, paragraph 0018). Examiner notes that the neural network is the trained proxy model and the verification dataset is the validation data set.);
Wan in view of Cimentada and Szeto are analogous to the claimed invention because they generate synthetic data to train machine learning models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan in view of Cimentada to verify the neural network using the verification dataset like in Szeto. Doing so “ensure[s] a proper analysis” (Szeto, page 24, paragraph 0121).
Wan in view of Cimentada, Szeto, and Hardy does not teach, but He does teach
and sorting the preliminary training dataset in descending order according to an importance of the features (He, page 5, paragraph 0009, “fault feature dimensionality reduction preprocessing, wherein an importance score of all features is calculated by ET algorithm, then and the features are sorted in descending order according a value of the importance score, and a new feature set is obtained by removing the features with a low importance score with a determined proportion.” Examiner notes that the features are the preliminary training dataset.).
Wan in view of Cimentada, Szeto, and Hardy and He are analogous to the claimed invention because they teach machine learning models that perform feature reduction. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan in view of Cimentada and Szeto to sort the training dataset in descending order like in Szeto. Doing so “avoids a long training time resulting from overly high dimensionality of the feature data” (He, page 10, paragraph 0069).
Claim(s) 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wan in view of Cimentada, Szeto, and Hardy, and Weng (“From GAN to WGAN”) (hereafter referred to as Weng).
Regarding claim 15, Wan in view of Cimentada, Szeto, and Hardy teach the method of Claim 1. Wan in view of Cimentada, Szeto, and Hardy does not teach, but Weng does teach
computing a Kullback-Leibler divergence between the original training dataset and the synthetic training dataset to determine a quality of the original training dataset (Weng, page 2, 2nd paragraph, “Before we start examining GANs closely, let us first review two metrics for quantifying the similarity between two probability distributions. (1) KL (Kullback-Leibler) divergence measures how one probability distribution p diverges from a second expected probability distribution q.” Examiner notes that p and q are the training dataset and synthetic training dataset and the quality of the training dataset is the similarity between p and q.).
Wan in view of Cimentada, Szeto, and Hardy and Weng are analogous to the claimed invention because they generate synthetic data to train prediction models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have implemented the prediction model in Wan in view of Cimentada and Szeto to use a Kullback-Leibler divergence like in Szeto. Doing so “quantif[ies] the similarity between two probability distributions” (Weng, page 2, 3rd paragraph).
Claim(s) 28 and 30-31 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wan in view of Cimentada, and in further view of Szeto, Hardy and Guillame-Bert et al. (US 2019/0251468 A1) (hereafter referred to as Guillame-Bert).
Regarding claim 28, Wan teaches
In response to all features of the original training dataset having been reproduced with synthetic data in the synthetic training dataset, append the synthetic training dataset to the original training dataset to form a hybrid training dataset having a larger size than the original training dataset; and training the machine learning model using the hybrid training dataset to increase an initial training accuracy of the machine learning model relative to training the machine learning model only using the original training dataset (Wan, page 3, 2nd paragraph, “On the last step, FFPred-GAN uses the Classifier Two-Sample Tests (CTST) to select the optimal synthetic training protein feature samples, which are used to augment the original training samples. During the down-stream machine learning classifier training stage, the optimal synthetic samples are expected to derive better classifiers, leading to higher predictive accuracy” where “the SVM trained by the augmented training protein feature samples learning those decision boundaries that successfully separate the protein samples distributed on the right corner of the figure” (Wan, page 9, last paragraph) and Wan, page 3, Figure 1
PNG
media_image1.png
338
748
media_image1.png
Greyscale
Examiner notes that the synthetic training dataset is the optimal synthetic training protein feature samples, the training dataset is the original training samples, and the hybrid training dataset is the augmented training protein feature samples. Examiner further notes that the augmenting the datasets is appending the synthetic training dataset to the training dataset to form a hybrid training dataset. Examiner also notes that the SVM is trained by the augmented training protein feature samples or the hybrid training dataset.) Examiner further notes that the hybrid dataset is shown in Figure 1 by combining the real and synthetic features into the critic which is larger than just the real feature samples alone. Examiner additionally notes that the higher predictive accuracy is an increase to an initial training accuracy.)
Wan does not teach, but Cimentada does teach
select a feature ci of the original training dataset as a target vector yi (Cimentada, page 1, 2nd paragraph, “Let’s imagine a data set with 30 rows. We separate the 1st row to be the test data.” Examiner notes that the first row is the target vector.);
select remaining features of the original training dataset as a set of training input vectors X\i, where X\i includes all features of the original training dataset other than a feature corresponding to the selected feature ci (Cimentada, page 1, 2nd paragraph, “We separate the 1st row to be the test data and the remaining 29 rows to be the training data.” Examiner notes that the training input vectors are the remaining 29 rows.);
train a prediction model f(yi|X\i) (Cimentada, page 1, 2nd paragraph, “We fit the model on the training data and then predict the one observation we left out.” Examiner notes that fitting the model is training a prediction model.);
generate an estimate y'i of the target vector yi by applying the prediction model to the set of training vectors X\i (Cimentada, page 1, 2nd paragraph, “We fit the model on the training data and then predict the one observation we left out.” Examiner notes that predicting the one observation left out is generating an estimate of the target vector.);
repeating, for a plurality of features of the original training dataset, operations of selecting a feature of the original training dataset, selecting remaining features of the original training dataset, training the prediction model, generating the estimate of the target vector (Cimentada, page 1, 2nd paragraph, “LOOCV: Let’s imagine a data set with 30 rows. We separate the 1st row to be the test data and the remaining 29 rows to be the training data. We fit the model on the training data and then predict the one observation we left out. We record the model accuracy and then repeat but predicting the 2nd row from training the model on row 1 and 3:30. We repeat until every row has been predicted.” Examiner notes that repeating until every row has been predicted is repeating for a plurality of features of the original training dataset, operations of selecting a feature of the original training dataset, selecting remaining features of the original training dataset, training the prediction mode, and generating the estimate of the target vector.)
Wan and Cimentada are considered analogous to the claimed invention because they both use the Leave One Out method on generated data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan to select features, train a model, and generate an estimate like in Cimentada. Doing so would be advantageous because “it uses all the data. At some point, every rows gets to be the test set and training set, maximizing information. In fact, it uses almost ALL the data as the original data set as the training set is just N-1” (Cimentada, page 2, bullet points).
Wan in view of Cimentada teach the estimate y'i of the target vector yi, but does not teach inserting a synthetic feature corresponding to the estimate into a synthetic training dataset. Szeto does teach
insert a synthetic feature c'i corresponding to the estimate y'i …into a synthetic training dataset (Szeto, page 9, paragraph 0018, “The modeling engine further generates one or more private data distributions from the local private data training set where the private data distributions represent the nature of the local private data used to create the trained model. The modeling engine uses the private data distributions to generate a set of proxy data, which can be considered synthetic data or Monte Carlo data having the same general data distribution characteristics as the local private data, while also lacking the actual private or restricted features of the local, private data….The modeling engine then attempts to validate that the set of proxy data is a reasonable training set stand-in for the local, private data by creating a trained proxy model from the set of proxy data” where “From the private data distributions, the machine learning engine can identify or otherwise calculate one or more salient private data features that describe the nature of the private data distributions” (Szeto, page 10, paragraph 0019) and where “private data servers transmit salient features of aggregated private data to a non-private computing devices which in turn creates proxy data for integrating into a trained global model” (Szeto, page 10, paragraph 0026). Examiner notes that the proxy data is the synthetic training dataset. Examiner further notes that the estimate is the salient private data features. By creating proxy data from the salient private data features, the synthetic features correspond to the estimate.)
repeating … inserting the synthetic feature into the synthetic training dataset (Szeto, page 24, paragraph 0118, “then the global modeling engine can repeat operations 660 through 680 until a satisfactory similar trained proxy model is generated” where “operation 660 shifts focus from the modeling engine in an entity’s private data server to the non-private computing device’s global modeling engine (see FIG. 1, global modeling engine 136). The global modeling engine receives the salient private data features and locally re-instantiates the private data features distributions in memory. As discussed previously with respect to operation 540, the global modeling engine generates proxy data from the salient private data features, for example, by using the re-instantiated private data distributions as probability distributions to generate new, synthetic sample data” (Szeto, page 23, paragraph 01115). Examiner notes that the repeating operation 660 is repeating inserting the synthetic feature into the synthetic training dataset.)
Wan in view of Cimentada and Szeto are analogous to the claimed invention because they generate synthetic data to train machine learning models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan in view of Cimentada to insert synthetic data into a dataset like in Szeto. Doing so “is considered advantageous because it provides for generating synthetic data capable of reproducing the knowledge gained from private data” (Szeto, page 18, paragraph 0078).
Wan in view of Cimentada and Szeto does not disclose, but Hardy does teach
transmitting, …, a message to at least one of the workers instructing the at least one worker to generate synthetic training data by instructing the at least one of the workers to: (Hardy, page 4, Figure 1b (see image below)
PNG
media_image2.png
400
626
media_image2.png
Greyscale
Examiner notes that the server sends the generator to the workers which maps to transmitting a message to the workers to generate synthetic training data.)
transmit weights of the machine learning model, trained using the hybrid training dataset, to the master node (Hardy, page 3, 2nd column, 1st paragraph, “Workers perform iterations locally on their data and every E epochs (i.e., each worker passes E times the data in their GAN) they send the resulting parameters to the server” where “The server generates a set K of k batches K = {X(1),…, X(k)}, with k ≤ N. Each X(i) is composed of b data generated by G. The server then selects, for each worker n, two distinct batches, say X(i) and X(j), which are sent to worker n and locally renamed as
X
n
(
g
)
and
X
n
(
d
)
. The way in which the two distinct batches are selected is discussed in Section IV-B1. Each worker n performs L learning iterations on its discriminator Dn (see Section II-1) using
X
n
(
d
)
and
X
n
(
r
)
, where
X
n
(
r
)
is a batch of real data extracted locally from Bn” (Hardy, page 3, 2nd column, first two bullet points) Examiner notes that performing iterations using both real and generated data is appending the synthetic training dataset to the training dataset to form a hybrid dataset and training the model. Examiner further notes that the indication from the master node is the server sending the generated data to the worker. Examiner additionally notes that the server is the master node and the workers transmit parameters or trained weights.)
and receiving, …, model parameters of a machine learning model from the at least one worker that were generated using the synthetic training data (Hardy, page 4, Figure 1b (see image below)
PNG
media_image2.png
400
626
media_image2.png
Greyscale
Examiner notes that since the discriminator is sent to the server and the discriminator has updated parameters, the model parameters from workers that were generated using the synthetic data were received by the server.)
wherein the model parameters received from the worker comprise trained neural network weights (Hardy, page 4, Figure 1b (see image below)
PNG
media_image2.png
400
626
media_image2.png
Greyscale
and “A GAN is a machine learning model, and more specifically a certain type of deep neural networks” (Hardy, page 1, 1st column, 3rd paragraph). Examiner notes that the updated parameters from the discriminator of the neural network are the trained neural network weights.).
Wan in view of Cimentada, Szeto and Hardy are analogous to the claimed invention because they generate synthetic data to train machine learning models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan in view of Cimentada and Szeto to use a federated learning system to generate synthetic data to train a model like in Hardy. Doing so “address[es] the problem of distributing GANs so that they are able to train over datasets that are spread on multiple workers” (Hardy, page 1, abstract).
Wan, Cimentada, Szeto, and Hardy do not explicitly disclose a message bus. Guillame-Bert, however does disclose
A method of operating a master node in a federated learning system including a plurality of workers that communicate with the master node via a message bus (Guillame-Bert, page 9, paragraph 0044, “DRF [Distributed Random Forest algorithm] computation can be distributed among computing machines called “workers”, and coordinated by a “manager”. The manager and the workers can communicate through a network.” Examiner notes that the network is the message bus.)
the message bus(Guillame-Bert, page 9, paragraph 0044, “DRF [Distributed Random Forest algorithm] computation can be distributed among computing machines called “workers”, and coordinated by a “manager”. The manager and the workers can communicate through a network.” Examiner notes that the network is the message bus.)
Wan, Cimentada, Szeto, Hardy and Guillame-Bert are analogous to the claimed invention because they use federated learning models. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to have implemented Wan, Cimentada, Szeto and Hardy on the federated learning system that includes the message bus in Guillame-Bert. Thus, this would be applying a known technique (federated learning) to a known device (message bus) ready for improvement to yield predictable results (communicate with master node) (MPEP 2143 I. (C) Use of known technique to improve similar devices (methods, or products) in the same way).
Regarding claim 30, Wan in view of Cimentada, Szeto, Hardy, and Guillame-Bert teach the method of Claim 28. Wan in view of Cimentada, Szeto, Hardy, and Guillame-Bert further teach
receiving from the at least one worker a set of preliminary neural network weights that were trained without using the synthetic training data (Hardy, page 5, 2nd column, 3rd paragraph, “Each worker n hosts a discriminator Dn and a training dataset Bn” where “Each discriminator n solely uses Bn- to train its parameters θn” (Hardy, page 5, 2nd column, 4th paragraph) and (Hardy, page 4, Figure 1b (see image below)
PNG
media_image2.png
400
626
media_image2.png
Greyscale
Examiner notes that the local training dataset Bn is used to train preliminary network weights or the parameters without using synthetic data since Bn is local to the discriminator. Examiner further notes that since the discriminator is sent to the server, the preliminary neural network weights or parameters are also received in the master node.);
and evaluating the set of preliminary neural network weights (Hardy, page 5, 2nd column, 3rd paragraph, “Each worker n hosts a discriminator Dn and a training dataset Bn” where “Each discriminator n solely uses Bn- to train its parameters θn” (Hardy, page 5, 2nd column, 4th paragraph) and “workers only have to handle their discriminator parameters θn and to compute error feedbacks after L local iterations” (Hardy, page 5, 2nd column, 5th paragraph). Examiner notes that computing error feedbacks based after local iterations is evaluating the set of preliminary neural network weights.),
wherein transmitting the message to the at least one worker instructing the at least one worker to generate the synthetic training data is performed in response to evaluating the set of preliminary neural network weights (Hardy, page 4, 1st column, 2nd to last paragraph, “the server generates new images to train all discriminators and updates w using error feedbacks” and where “workers only have to handle their discriminator parameters θn and to compute error feedbacks after L local iterations” (Hardy, page 5, 2nd column, 5th paragraph). Examiner notes that the server or master node generates synthetic data using error feedbacks sent from the worker where the computing the error feedback is evaluating the set of preliminary neural network weights.).
Wan in view of Cimentada, Szeto and Hardy are analogous to the claimed invention because they generate synthetic data to train machine learning models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan in view of Cimentada and Szeto to use a federated learning system to generate synthetic data to train a model like in Hardy. Doing so “address[es] the problem of distributing GANs so that they are able to train over datasets that are spread on multiple workers” (Hardy, page 1, abstract).
Regarding claim 31, Wan in view of Cimentada, Szeto, Hardy, and Guillame-Bert teach the method of Claim 28. Wan in view of Cimentada, Szeto, Hardy, and Guillame-Bert further teach
after instructing the at least one worker to generate the synthetic training data, receiving a quality metric from the at least one worker, wherein the quality metric measures a quality of a synthetic training dataset (Hardy, page 4, 2nd column, 2nd paragraph, “Every global iteration, the server receives the error feedback Fn from every worker n, corresponding to the error made by G on
X
n
(
g
)
” where “The server generates a set K of k batches K = {X(1),…, X(k)}, with k ≤ N. Each X(i) is composed of b data generated by G. The server then selects, for each worker n, two distinct batches, say X(i) and X(j), which are sent to worker n and locally renamed as
X
n
(
g
)
and
X
n
(
d
)
” (Hardy, page 3, 2nd column, first bullet point) Examiner notes that G is the generator on the master node or server,
X
n
(
g
)
is the generated or synthetic training dataset, and the error feedback is the quality metric. Since the error feedback is computed on the synthetic dataset from the server, receiving the quality metric happens after instructing the worker to generate synthetic training data.);
and instructing the worker to proceed with training a machine learning model using the synthetic training dataset in response to the quality metric (Hardy , page 5, Algorithm 1 (see image below)
PNG
media_image3.png
848
413
media_image3.png
Greyscale
Examiner notes that between lines 27 and 40, the server iteratively sends generated data (line 34) and gets feedback about the generated data (line 36). After computing gradients and updating the gradients (lines 37-39), the server proceeds to start the for loop over in which the server once again sends generated data to the worker. Examiner further notes that between lines 2 and 13, the worker also iteratively receives the generated data from the server (line 5) and sends feedback to the server (line 10). By iteratively performing the steps of sending data and receiving feedback, the server is instructing the worker to proceed with training.).
Wan in view of Cimentada, Szeto and Hardy are analogous to the claimed invention because they generate synthetic data to train machine learning models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Wan in view of Cimentada and Szeto to use a federated learning system to generate synthetic data to train a model like in Hardy. Doing so “address[es] the problem of distributing GANs so that they are able to train over datasets that are spread on multiple workers” (Hardy, page 1, abstract).
Claim(s) 45 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wan in view of Cimentada, Szeto, and Hardy, and Chen et al. (“LAG: Lazily Aggregated Gradient for Communication-Efficient Distributed Learning”) (hereafter referred to as Chen).
Regarding claim 45, Wan in view of Cimentada, Szeto, and Hardy teach the method of Claim 1. Wan in view of Cimentada, Szeto, and Hardy do not explicitly teach, but Chen does teach
wherein the increased initial training accuracy reduces a number of iterations of communication with the master node needed to train the machine learning model to a threshold accuracy, thereby reducing a network footprint of a federated learning system (Chen, page 1, abstract, “Theoretically, the merits of this contribution are: i) the convergence rate is the same as batch gradient descent in strongly convex, convex, and nonconvex cases; and, ii) if the distributed datasets are heterogeneous (quantified by certain measurable constants), the communication rounds needed to achieve a targeted accuracy are reduced thanks to the adaptive reuse of lagged gradients.” Examiner notes that the communication rounds are the iterations.)
Wan in view of Cimentada, Szeto, and Hardy and Chen are analogous to the claimed invention because they teach distributed networks. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have implemented the prediction model in Wan in view of Cimentada, Szeto, and Hardy to reduce the number of iterations of communication. Doing so is advantageous because “the communication rounds needed to achieve a targeted accuracy are reduced thanks to the adaptive reuse of lagged gradients” (Chen, page 1, abstract).
Response to Arguments
Applicant’s arguments with respect to 101 on pages 13-16 and in light of the instant
amendments have been fully considered and are persuasive. Specifically, the argument on pages
14-15 regarding the improvement recited in the specification and reflected in the claims was persuasive. The 101 rejections of the claims have been withdrawn. Examiner notes that the claims reflect the improvement of improving model accuracy, decreasing training time, and/or reducing the network footprint needed specifically with the newly amended limitations of “form a hybrid training dataset having a larger size than the original dataset” and “training the machine learning model using the hybrid training dataset to increase an initial training accuracy of the machine learning model relative to training the machine learning model only using the original training dataset” in light of the claim as a whole.
On page 11, Applicant argues:
The Office has maintained that Wan's "augmenting" teaches "appending the synthetic training dataset to the original training dataset to form a hybrid training dataset." Although Applicant respectfully disagrees, the advance prosecution, claim 1 has been amended to further clarify that, "in response to all features of the original training dataset having been reproduced with synthetic data in the synthetic training dataset, appending the synthetic training dataset to the original training dataset to form a hybrid training dataset having a larger size than the original training dataset." However, Wan does not append a synthetic dataset to an original dataset to form a new, larger hybrid dataset. Rather, Wan "select[s] the optimal synthetic training protein feature samples, which are used to augment the original training samples." (See, page 3, 2nd paragraph of Wan). Therefore, Wan selects a curated subset of synthetic samples, via Classifier Two-Sample Tests, and uses that subset to modify the original training samples. There is no disclosure in Wan of "all features of the original training dataset having been reproduced with synthetic data in the synthetic training dataset, appending the synthetic training dataset to the original training dataset to form a hybrid training dataset having a larger size than the original training dataset," as recited by claim 1. In other words, Wan's selection of an optimal subset of synthetic samples is not the generation of a complete synthetic dataset in which all features have been reproduced and then used to create a hybrid training dataset "having a larger size than the original training dataset." As such, Wan cannot reasonably be construed to teach or suggest the above-noted features of claim 1.
Regarding the Applicant’s argument that Wan does not teach the limitation of “in response to all features of the original training dataset having been reproduced with synthetic data in the synthetic training dataset, appending the synthetic training dataset to the original training dataset to form a hybrid training dataset having a larger size than the original,” Examiner respectfully disagrees. Specifically, Examiner notes that Wan does teach reproducing original data with synthetic data (Wan, page 3, Figure 1). Examiner notes that in Figure 1, the original and synthetic data are both entered into the critic. Examiner notes the broadest reasonable interpretation of “reproduced with” includes producing the synthetic and original data to be used together as Figure 1 is depicting. Examiner further notes that Figure 1 further shows that combining the original and synthetic data creates a dataset that is larger than the original training dataset.
On pages 11-13, Applicant argues:
Furthermore, Wan, alone or in combination with Cimentada does not teach or suggest the claimed column-wise, prediction-based synthetic-feature generation. The claims recite a method that operates on features (columns). As an example, a feature is selected as the target vector, Ci the remaining features serve as training inputs, a prediction model is trained to predict the selected feature from the remaining features, and the resulting estimate is inserted as a synthetic feature into a synthetic training dataset. This is a data-generation technique that produces new synthetic feature values. Moreover, based on this Amendment, this process be repeated for a plurality of features. Thereafter, the synthetic training dataset is appended to the original training dataset "in response to all features of the original training dataset having been reproduced with synthetic data in the synthetic training dataset." Thus, the claims provide steps in which a complete synthetic dataset is generated, column by column, before the synthetic dataset is appended to the original dataset. Cimentada, by contrast, teaches leave-one-out cross-validation ("LOOCV"), which operates on observations (rows) and is performed to predict values.
As the Office's own reproduction of Cimentada confirms, "Let's imagine a data set with 30 rows. We separate the 1st row to be the test data and the remaining 29 rows to be the training data. We fit the model on the training data and then predict the one observation we left out. We record the model accuracy and then repeat but predicting the 2nd row from training the model on row 1 and 3:30. We repeat until every row has been predicted." Cimentada, page 1, 2nd paragraph. Therefore, Cimentada leaves out a row (an observation/data point) to measure model accuracy, and does not provide a feature (column) to generate a synthetic feature that is inserted into a synthetic training dataset. In fact, Cimentada does not produce any synthetic data, it produces accuracy scores.
Moreover, even if the predictions of Cimentada could be reasonably be construed to teach the generation of synthetic data, Cimentada fails to teach or suggest that all of the "all features of the original training dataset having been reproduced with synthetic data in the synthetic training dataset", much less that the results of the process performed by Cimentada "form a hybrid training dataset having a larger size than the original training dataset." Lastly, the Office's assertion that "the first row is the target vector" conflates rows (observations) with features (columns), which are fundamentally different dimensions of a dataset as understood by one of ordinary skill in the art. As such, Cimentada cannot reasonably be construed to teach or suggest the above-noted features of claim 1.
Szeto does not cure this deficiency. Instead, Szeto discusses how "The modeling engine uses the private data distributions to generate a set of proxy data, which can be considered synthetic data or Monte Carlo data having the same general data distribution characteristics as the local private data."). (See Szeto, paragraph [0018]). Generating data by sampling from probability distributions is a fundamentally different technique from the claimed approach of predicting one feature from the remaining features using a trained prediction model and inserting the resulting estimate as a synthetic feature. As such, Szeto cannot reasonably be construed to teach or suggest the above-noted features of claim 1.
Regarding the Applicant’s argument that Wan in combination does not disclose the column-wise, prediction-based synthetic-feature generation, Examiner respectfully disagrees. Specifically, Examiner notes that claim 1 recites features, not columns. Thus, Wan in view of Cimentada, Szeto and Hardy do not have to teach column-wise prediction. Examiner further notes that as discussed in the 112(b) rejection above, claim 1 does not recite producing original data and thus, it is unclear how this original data can be reproduced with synthetic data. As such, Examiner has interpreted this specific limitation of “in response to all features of the original training dataset having been reproduced with synthetic data in the synthetic training dataset” to be generating synthetic training data using original training data. Examiner further notes that as discussed in the prior argument above, Wan teaches this as well as the limitation of “form a hybrid training dataset having a larger size than the original training dataset”.
On page 13, Applicant argues:
Dependent Claims 2, 4-5, 7, 9-12, 15, 30-31, and 37-42 are Patentable
The dependent claims are patentable at least per the patentability of the independent claims from which they depend. Applicant traverses the rejections of the remaining dependent claims. However, as each of these claims depends from a base claim that is believed to be in condition for allowance, Applicant does not believe that it is necessary to argue the allowability of each of the remaining dependent claim individually. Applicant does not necessarily concur with the interpretation of these claims or with the bases for rejection set forth in the Office Action. Applicant therefore reserves the right to address the patentability of these claims individually as necessary in the future.
New Claims
Claims 45 and 46 are added by this Amendment and are believed to be allowable for the following reasons. Claims 45 and 46 depend from claim 1 and are distinguishable from the applied art for at least the same reasons as the base claim. Moreover, Claims 45 and 46 recite additional combinations of features that are not taught by the applied art.
Regarding the Applicant’s argument that the dependent claims are allowable at least due in part to their dependency on the independent claims, the Examiner respectfully disagrees and notes the instant rejections and response to arguments regarding the independent claims above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Zhang et al. (“Poisoning Attack in Federated Learning using Generative Adversarial Nets”) also discusses using a GAN in a federated learning environment.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KAITLYN R LAU whose telephone number is (571)272-1429. The examiner can normally be reached Monday - Thursday: 8:00 am - 6:00 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/K.R.L./Examiner, Art Unit 2148
/MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148