Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim(s) 1-20 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
As per claim 1, line 13-14, it is unclear what constitutes a “mixture distribution from among the tuned at least one synthesizer”. It is unclear how there can be mixture distribution formed from one single synthesizer. For the purpose of examination, the limitation will be interpreted as forming a mixture distribution from a plurality of synthesizers, and generating mixed variable output from tuned synthesizers.
Claims 10 and 19 recite similar limitations as claim 1 and therefore are rejected under the same rationale as claim 1. Claims 2-9, 11-18, and 20 are dependent on Claims 1, 10, and 19 respectively and do not add enough to overcome this rejection. Therefore, they are rejected under the same rationale.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 5-10, 14-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kar et al. (Meta-Sim: Learning to Generate Synthetic Datasets, hereinafter Kar) in view of Breugel et al. (Synthetic Data, Real Errors: How (Not) to Publish and Use Synthetic Data, hereinafter Breugel)
Regarding Claim 1:
Kar discloses:
A method for facilitating supervised generative optimization for synthetic data generation, the method being implemented by at least one processor ([Page 4557 right col, second para] discloses utilizing a Titan Xp GPU), the method comprising: ([Page 4553 left col first para; page 4553 right col third para; page 4555 Section 3.2.3; Page 4557 Section 4.3 right col last para; Fig 4] Page 4553 right column, paragraph 3, discloses using a distribution transformer to sample scenes and generate synthetic datasets. (i.e. synthetic data). Page 4555, section 3.2.3 and Fig 4 discloses optimizing the generation of synthetic data by optimizing downstream tasks. Page 4557, section 4.3 right column last paragraph, discloses the KITTI dataset that comes with labels (i.e. supervised).)
receiving, by the at least one processor via an application programming interface, at least one input, each of the at least one input including input data and at least one parameter; ([Page 4555 right col Algorithm 1; Page 4557 right col second para; Page 4556 right col second para; Page 4552 right col last para; page 4555 lefto col 2nd para] discloses Given values to the algorithm that represents the optimization process. XR represents real images (i.e. input data) while the other given value P represents probabilistic grammar. This probabilistic grammar is additional input data that uses random parameters (i.e. parameters.) as disclosed on page 4556 right column paragraph 2. Page 4552, right column, last paragraph discloses parameter θ. Parameter θ is further disclosed on page 4555, left column, second paragraph. Page 4557 right column, paragraph 2 discloses utilizing a GPU, and Algorithm 1 on Page 4555 discloses the given variable XR, which are real images. It is obvious that the real image data would need to be received from an input interface, with the help of a GPU.)
partitioning, by the at least one processor, the input data to generate at least one data set, the at least one data set including at least one from among a training data set, a validation data set, and a test data set; ([Page 4557 Section 4.3 right col last para] beginning with “Experimental Setup”, discloses a validation data formed from the inputted “KITTI” dataset. The paragraph continues onto the left column, explaining that 100 images from the “KITTI” dataset is taken to form validation set “V” and the rest form the real dataset “XR”.)
tuning, by the at least one, at least one hyperparameter of at least one synthesizer by using the at least one data set and supervised optimization that is based on at least one downstream performance metric; ([Page 4555 Algorithm 1; Page 4552 left col para 1] discloses utilizing hyperparameters in a loop to generate synthetic images. These hyperparameters include Epochs, Iters, and Batch Size. The hyperparameters in the loop are used in image generation, which are then used to calculate a loss based on the generation. The loop allows for tuning by allowing the process to change hyperparameter values of generator “G0” (i.e. Synthesizer) until the generation meets a loss function value calculate using validation set “V”. Page 4552, left column paragraph 1, discloses optimizing a meta objective to improve a downstream performance of a task model trained by synthetic datasets. (i.e tuning based on downstream performance and at least one data set). In addition, the generated images are rendered using labels from the samples, “S”, of probabilistic grammar “P” (i.e. supervised))
Kar does not explicitly disclose: determining, by the at least one processor, a mixture distribution from among the tuned at least one synthesizer;
training, by the at least one processor, at least one model based on the mixture distribution;
and generating, by the at least one processor using the trained at least one model, at least one set of synthetic data based on the input data.
However, Breugel discloses in the same field of endeavor: determining, by the at least one processor, a mixture distribution from among the tuned at least one synthesizer; ([Page 13 Section A.2 Sections 4.1; Fig 1; Page 4 Section 3.4] Figure 1 shows an example of how the deep generative ensemble generates a possible distribution of the real dataset (i.e. mixture distribution). Section 3.4 on Page 4 discloses approximating the distribution of the real dataset as p and obtaining it by generating synthetic data sets k numerous amounts of times from a generative model (i.e. synthesizer))
training, by the at least one processor, at least one model based on the mixture distribution; ([Page 13 Section A.2 Section 4.1; Page 4 Section 4.1 right col para 4] discloses training downstream approaches based on the ensemble of synthetic datasets. These downstream approaches include training models. Page 4, section 4.1, paragraph 4, right col discloses training a model on each set of synthetic data from the distribution, as shown in equation two. (i.e. training based on a mixture distribution))
and generating, by the at least one processor using the trained at least one model, at least one set of synthetic data based on the input data. ([Page 13 both col section 4.1 pipeline] discloses generating synthetic data by a model in step 3 of the pipeline, after the generative model is trained. Steps 2-5 are repeated at step 6 to generate synthetic data based on the input data after training occurs in step 2.)
Kar and Breugel are both analogous art to the present invention because both are from the same field of endeavor directed to generating synthetic data and optimizing the quality of synthetic data generation.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the Meta-Sim experiment of generating synthetic datasets, and the similar approach to generating synthetic data in an ensemble by Breugel. One would be motivated to add the feature of the ensemble model in generating synthetic datasets in order to achieve a more accurate synthesized training dataset from a distribution of synthesizers, as shown by downstream performance metrics. This would net a more rigorously tested dataset than simply generating a synthetic dataset from a single operation of training.
Regarding Claim 5:
Kar in view of Breugel discloses: The method of claim 1, as discussed above.
Breugel further discloses: wherein the mixture distribution corresponds to an automatically determined composition of at least one variable that is derived from output of each of the tuned at least one synthesizer, the at least one variable including a random variable. ([Page 3 left col para 2] starting with section 3.1. Set-up, discloses random variables from a distribution outputted from generator G0 (i.e. tuned synthesizer). The generators are utilized in a deep generative ensemble to create a mixture distribution as disclosed in claim 1. In addition, variable T is described in the section as a random variable in relation to the distribution.)
Regarding Claim 6:
Kar in view of Breugel discloses: The method of claim 1, as discussed above.
Kar further discloses: wherein the tuning of the at least one hyperparameter of the at least one synthesizer comprises:
optimizing, by the at least one processor, at least one target function that corresponds to each of the at least one synthesizer, ([Page 4555 Algorithm 1] discloses tuning hyperparameters as discussed above. The same section discloses optimizing G0 (i.e. synthesizer) using a loss function (i.e. target function) which is achieved by utilizing a validation set of data.)
wherein the optimizing relates to a bi-level optimization of the at least one target function. ([Page 4555 Algorithm 1] discloses nested loop (i.e. bi-level) optimization of a loss function. The inner loop calculates loss based on a training loss (based on the real images) and the outer loop augments the loss based on a validation loss (based on a target validation set))
Regarding Claim 7:
Kar in view of Breugel discloses: The method of claim 6, as discussed above.
Kar further discloses: wherein the optimizing of the at least one target function further comprises:
minimizing, by the at least one processor, at least one validation loss function that corresponds to each of the at least one synthesizer based on the validation data set; ([Page 4555 Algorithm 1; Page 4555 left col para 3] The algorithm discloses calculating and optimizing a loss function of G0 (i.e. synthesizer) which is derived from a score that is calculated using the validation dataset. Paragraph 3 on the left column of the same page discloses reducing variance of the loss estimator to bolster the loss representation.)
and minimizing, by the at least one processor, at least one training loss function that corresponds to each of the at least one synthesizer based on the training data set. ([Page 4555 Algorithm1] discloses calculated loss in the inner loop utilizing the training data set XR and MMD (difference calculations between distributions). This loss is used in the optimize function outside the loop, alongside the validation loss, to minimize the loss. The loss is minimized by a SGD step, which adjusts hyperparameters of the model to optimize the loss.)
Regarding Claim 8:
Kar in view of Breugel discloses: The method of claim 1, as discussed above.
Breugel further discloses: further comprising:
Augmenting, by the at least one processor, the input data by incorporating the generated at least one set of synthetic data into the input data, ([Page 8 left col para 2] discloses how the DGE approach is a data augmentation method, where the generated synthetic data sets are incorporated into the input data.)
Wherein the input data is augmented for the at least one downstream performance metric for a corresponding downstream task. ([Page 7 right col para 3-5; Page 4 right col para 4; Page 8 left col para 1] Page 7 discloses using a naïve model and a DGE model trained using the COVID 19 data set to measure the performance between using a model trained on a distribution of synthetic data as opposed to a singular set of synthetic data. This is further stated on Page 4, right column, paragraph 4, where the experiment is shown to compare a singular synthetic dataset vs a distribution of synthetic datasets. Page 8, left column, paragraph 1 discloses how the performance of the downstream model is affected and improved by the augmentation of data done by the deep generative ensemble.)
Regarding Claim 9:
Kar in view of Breugel discloses: The method of claim 1, as discussed above.
Breugel further discloses: wherein each of the at least one synthesizer and the at least one model includes at least one from among a deep learning model, a neural network model, a machine learning model, a mathematical model, and a process mode. ([Page 13 right col para 3] discloses using a CTGAN (i.e. deep learning model) to generate synthetic data (i.e. synthesizer). It is additionally disclosed that the downstream model (i.e. at least one model) is an MLP model (i.e. neural network model).)
Regarding Claim 10:
Claims 10 and 19 recite a machine (system) and an article of manufacture that performs the same method as described in Claim 1. Therefore, claims 10 and 19 are rejected under the same reasons mentioned for Claim 1. The additional elements of claims 10 and 19 are addressed below by Breugel:
Claim 10: A computing device configured to implement an execution of a method for facilitating supervised generative optimization for synthetic data generation, the computing device comprising: a processor; a memory; and a communication interface coupled to each of the processor and the memory ([Page 16 right col para 1] discloses training the CTGAN on a consumer PC, including an RTX 3080 gpu (i.e. processor, memory, interface))
Regarding Claims 14:
Claim 14 recites a machine (system) that performs the method as described in claim 5. Therefore, claim 14 is rejected under the same reasons mentioned for claim 5. The claim’s additional elements are disclosed by Breugel as shown in the rejection of claim 10, as claim 14 is dependent on claim 10.
Additional elements: The computing device of claim 10 ([Page 16 right col para 1] discloses training the CTGAN on a consumer PC, including an RTX 3080 gpu (i.e. processor, memory, interface))
Regarding Claims 15-18:
Claims 15-18 recites a machine (system) that performs the method as described in claims 6-9. Therefore, claims 15-18 are rejected under the same reasons mentioned for claims 6-9. The claims’ additional elements are disclosed by Breugel as shown in the rejection of claim 10, as claims 15-18 are dependent on claim 10.
Additional elements: The computing device of claim 10 ([Page 16 right col para 1] discloses training the CTGAN on a consumer PC, including an RTX 3080 gpu (i.e. processor, memory, interface))
Regarding Claim 19:
Claim 19 recites an article of manufacture that performs the same method as described in Claim 1. Therefore, claim 19 is rejected under the same reasons mentioned for Claim 1. The additional elements of claim 19 is addressed below by Breugel:
A non-transitory computer readable storage medium storing instructions ([Page 16 right col para 1] discloses training the CTGAN on a consumer PC, including an RTX 3080 gpu. A consumer PC typically will include storage devices such as hard drives (i.e. non-transitory computer readable storage medium.))
Claim(s) 2, 11, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kar et al. (Meta-Sim: Learning to Generate Synthetic Datasets, hereinafter Kar) in view of Breugel et al. (Synthetic Data, Real Errors: How (Not) to Publish and Use Synthetic Data, hereinafter Breugel) in view of Okada et al. (US 20200257824 A1, hereinafter Okada)
Regarding Claim 2:
Kar discloses in view of Breugel: The method of claim 1 as discussed above.
Kar in view of Bruegel does not explicitly disclose: wherein the input data includes at least one collection of tabular data for synthesis according to the at least one parameter, the at least one parameter including a number of tabular rows for the synthesis.
However, Okada discloses in the same field of endeavor: The method of claim 1, wherein the input data includes at least one collection of tabular data for synthesis according to the at least one parameter, the at least one parameter including a number of tabular rows for the synthesis. ([Para 2; Para 21] Page 1, paragraph 2 discloses that tabular data in the synthetic data generation apparatus is in the format of a matrix, with rows being called a record. On page 2, paragraph 21, it is disclosed that the input takes tabular data and records n in as input (i.e. tabular data and number of rows).)
Kar, Breugel, and Okada are analogous art to the present invention because all three are from the same field of endeavor directed to generating synthetic data.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the Meta-Sim experiment of generating synthetic datasets, the similar approach to generating synthetic data in an ensemble by Breugel, and the method for generating synthetic data based on numerical attributes from tabular data format. One would be motivated to add the feature of generating data from tabular formats because it gives a wider range of different types of data to synthesize and optimize to the present invention.
Regarding Claim 11:
Claim 11 recites a machine (system) that performs the same method as described in Claim 2, in addition to the additional elements of claim 10 of which claim 11 is dependent on. Therefore, claim 11 is rejected under the same reasons mentioned for Claim 2.
The additional elements of claim 10 are rejected earlier by Kar in view of Breugel in an earlier rejection.
Regarding Claim 20:
Claim 20 recites an article of manufacture that performs the same method as described in Claim 2, in addition to the additional elements of claim 19 of which claim 20 is dependent on. Therefore, claim 20 is rejected under the same reasons mention for Claim 2.
Additional elements: The storage medium of claim 19 ([Page 16 right col para 1] discloses training the CTGAN on a consumer PC, including an RTX 3080 gpu. A consumer PC typically will include storage devices such as hard drives (i.e. non-transitory computer readable storage medium.))
Claim(s) 3 and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kar et al. (Meta-Sim: Learning to Generate Synthetic Datasets, hereinafter Kar) in view of Breugel et al. (Synthetic Data, Real Errors: How (Not) to Publish and Use Synthetic Data, hereinafter Breugel) in view of Franceschi et al. (Bilevel Programming for Hyperparameter Optimization and Meta-Learning, hereinafter Franceschi).
Regarding Claim 3:
Kar in view of Breugel discloses: The method of claim 1, as discussed above.
Kar further discloses: wherein the at least one hyperparameter is tuned based on at least one downstream performance metric ([Page 4555 Algorithm 1] discloses a Task network that determines a score and a loss (i.e. downstream performance metric) which are dependent on hyperparameters epochs, iterations, and batch size, all in a loop to continuously tune said hyperparameters.)
Kar in view of Breugel does not disclose:
However, Franceschi discloses in the same field of endeavor: ([Page 2 right col para 3] beginning with “2.1. Hyperparameter Optimization”, discloses on page 2, right column, paragraph 3, optimizing hyperparameters with the goal of minimizing the validation error (i.e. regularization in terms of the spec) of a model through the use of regularization hyperparameters to control space (i.e. regularizing fidelity) or penalties. The section further discloses parameter “Ω” (omega) to improve the fit of the data (i.e. fidelity).)
Kar, Breugel, and Franceschi are analogous art to the present invention because all three are from the same field of endeavor directed to generating synthetic data and optimizing the process of generating synthetic data.
It would be obvious for on of ordinary skill in the art before the effective filing date of the claimed invention to have combined the Meta-Sim experiment of generating synthetic datasets, the similar approach to generating synthetic data in an ensemble by Breugel, and the method of hyperparameter optimization in meta learning by Franceschi. One would be motivated to add regularization parameters in hyperparameter tuning to allow for a better fit of synthetic data generated with respect to the real dataset in synthetic data generation.
Regarding Claim 12:
Claim 12 recites a machine (system) that performs the method as described in Claim 3. Therefore, Claim 12 is rejected under the same reasons mentioned for claim 3. Claim 12’s additional elements are disclosed by Breugel as shown in the rejection of claim 10, as claim 12 is dependent on claim 10.
Additional elements: The computing device of claim 10 ([Page 16 right col para 1] discloses training the CTGAN on a consumer PC, including an RTX 3080 gpu (i.e. processor, memory, interface))
Claim(s) 4 and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kar et al. (Meta-Sim: Learning to Generate Synthetic Datasets, hereinafter Kar) in view of Breugel et al. (Synthetic Data, Real Errors: How (Not) to Publish and Use Synthetic Data, hereinafter Breugel) in view of Wagh et al. (US 20240054405 A1, hereinafter Wagh)
Regarding Claim 4:
Kar in view of Breugel discloses: The method of claim 1 as discussed above.
Kar in view of Bruegel does not explicitly disclose: wherein each of the at least one synthesizer corresponds to a synthetic data generator that uses a synthetic data generation algorithm to identify at least one property of sampled data, the at least one property including at least one from among a correlation, a distribution, and a pattern.
However, Wagh discloses in the same field of endeavor: wherein each of the at least one synthesizer corresponds to a synthetic data generator that uses a synthetic data generation algorithm to identify at least one property of sampled data, the at least one property including at least one from among a correlation, a distribution, and a pattern. ([Para 57; Para 83] Page 6, paragraph 57 discloses satellite systems (i.e. synthesizer) which include a synthetic data generation pipeline. Page 10, paragraph 83 discloses the satellite systems utilizing synthetic data generation algorithms to process the input data as a distribution and create datapoints for the distribution. (i.e. identify a distribution))
Kar, Breugel, and Wagh are analogous art to the present invention because all three are from the same field of endeavor directed to generating synthetic data.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the Meta-Sim experiment of generating synthetic datasets, the similar approach to generating synthetic data in an ensemble by Breugel, and the method for federated machine learning proposed by Wagh, particularly focusing on synthetic data generation. One would be motivated to add synthetic data generators using synthetic data generation algorithms to allow for a better representation of input data by identify distributions of said input data.
Regarding Claims 13:
Claim 13 recites a machine (system) that performs the method as described in claim 4. Therefore, claim 13 is rejected under the same reasons mentioned for claim 4. The claim’s additional elements are disclosed by Breugel as shown in the rejection of claim 10, as claim 13 is dependent on claim 10. ([Page 16 right col para 1] discloses training the CTGAN on a consumer PC, including an RTX 3080 gpu (i.e. processor, memory, interface))
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SUMAIR R CHOWDHURY whose telephone number is (571)270-0523. The examiner can normally be reached Monday-Friday 8:30am - 5:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ABDULLAH AL KAWSAR can be reached at (571) 270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SUMAIR RASHED CHOWDHURY/Examiner, Art Unit 2127 7/30/2026
/TEWODROS E MENGISTU/Examiner, Art Unit 2127