Prosecution Insights
Last updated: August 17, 2026
Application No. 19/295,869

METHOD AND APPARATUS FOR GENERATING DATA

Non-Final OA §101
Filed
Aug 11, 2025
Priority
Sep 23, 2024 — RE 10-2024-0128405
Examiner
LE, HUNG D
Art Unit
2161
Tech Center
2100 — Computer Architecture & Software
Assignee
Electronics and Telecommunications Research Institute
OA Round
1 (Non-Final)
90%
Grant Probability
Favorable
1-2
OA Rounds
1y 4m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 90% — above average
90%
Career Allowance Rate
984 granted / 1092 resolved
+35.1% vs TC avg
Moderate +6% lift
Without
With
+6.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 4m
Avg Prosecution
19 currently pending
Career history
1115
Total Applications
across all art units

Statute-Specific Performance

§101
14.0%
-26.0% vs TC avg
§103
41.3%
+1.3% vs TC avg
§102
20.5%
-19.5% vs TC avg
§112
8.2%
-31.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1092 resolved cases

Office Action

§101
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION 1. This Office Action is in response to the application filed on 08/11/2025. Claims 1-20 are pending. Priority 2. Receipt is acknowledged of papers submitted under 35 U.S.C. 119(a)-(d), which papers have been placed of record in the file. Information Disclosure Statement 3. The information disclosure statement (IDS) filed on 08/11/2025 complies with the provisions of M.P.E.P. 609. The examiner has considered it. Claim Rejections - 35 USC § 101 4. 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. 5. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. At Step 1: Independent claims 1, 13 and 20 are directed to a “method” and an "apparatus” and thus directed to a statutory category At Step 2A, Prong One: The claim recites the following limitations directed to an abstract idea: • " receiving original data and a desired data quality level from a user" as drafted this recites a mentally performable process as an evaluation or judgement. This is also consistent with the specification as in Fig. 9 and paragraph 87 where one can mentally visualize receiving various data from a user. • " determining an original data quality level of the original data " as drafted this recites a mentally performable process as an evaluation or judgement. This is also consistent with the specification as in Fig. 9 and paragraph 87 where one can mentally visualize receiving various data from a user. • " determining a data generation model by additionally training an initial model, which is pre-trained to receive an input data quality level as input and generate output data of the input data quality level, using the original data and the original data quality level " as drafted this recites a mentally performable process as an evaluation or judgement. This is also consistent with the specification as in Fig. 9 and paragraph 88 where one can mentally visualize training or pre-training model to process incoming data. • " generating synthetic data by executing the data generation model using the desired data quality level" as drafted this recites a mentally performable process as an evaluation or judgement. This is also consistent with the specification as in Fig. 9 and paragraph 89 where one can mentally visualize generating synthetic data using a data generation model. • " generating merged data by combining the original data with the synthetic data" as drafted this recites a mentally performable process as an evaluation or judgement. This is also consistent with the specification as in Fig. 9 and paragraph 90 where one can mentally visualize merging data. • " providing the merged data to the user" as drafted this recites a mentally performable process as an evaluation or judgement. This is also consistent with the specification as in Fig. 9 and paragraph 91 where one can mentally visualize providing merged data to the user. At Step 2A, Prong Two: • The claim recites no additional elements. At most one might consider that a "circuitry configured to" as claimed might be considered to represent a computer-implemented system and method consistent with Fig. 1 even though the claim does not recite any computer. At most this would be a high-level recitation of a generic computer components and represents mere instructions to apply the abstract idea on a computer as in MPEP 2106.05(f), which does not provide integration into a practical application. • Viewing the additional limitations together and the claim as a whole, nothing provides integration into a practical application. At Step 2B: • The conclusions for the mere implementation using a computer are carried over and does not provide significantly more. • Looking at the claim as a whole does not change this conclusion and the claim is ineligible. Dependent Claims 2-12 and 14-19 The limitations as recited in dependent claims 2 and 14 recite, “determining a synthetic data quality target ……” which further describes the concepts performed in the human mind including an observation, evaluation, judgment, and opinion, in step 2A prong one. The limitations as recited in dependent claims 3 and 15 recite, “… evaluating whether the synthetic data satisfied the synthetic data quality target ………” which further describes the concepts performed in the human mind including an observation, evaluation, judgment, and opinion, in step 2A prong one. The limitations as recited in dependent claims 4 and 16 recite, “… evaluating whether the merged data satisfied the desired data quality level ………” which further describes the concepts performed in the human mind including an observation, evaluation, judgment, and opinion, in step 2A prong one. The limitations as recited in dependent claims 5 and 17 recite, “… determining a merged rule for selecting data to be combined with the original data … generating the merged data ………” which further describes the concepts performed in the human mind including an observation, evaluation, judgment, and opinion, in step 2A prong one. The limitations as recited in dependent claims 6 and 17 recite, “…evaluating whether the merged data satisfied the desired data quality level ………” which further describes the concepts performed in the human mind including an observation, evaluation, judgment, and opinion, in step 2A prong one. The limitations as recited in dependent claim 7 recites, “…generating new synthetic data … generating the new merged data………” which further describes the concepts performed in the human mind including an observation, evaluation, judgment, and opinion, in step 2A prong one. The limitations as recited in dependent claim 8 recites, “…predicting whether the merged data satisfies the desired data quality level … generating new synthetic data …generating the merged dasta………” which further describes the concepts performed in the human mind including an observation, evaluation, judgment, and opinion, in step 2A prong one. The limitations as recited in dependent claim 9 recites, “wherein the desired data quality level comprises one or more data evaluation factors...” which further describes the concept is mere gathered data under prong 2 (insignificant extra solution activity— MPEP 2106.06g) and WURC under 2b (using gather data - MPEP 2106.05d). The limitations as recited in dependent claim 10 recites, “wherein the data generation model is determined by additionally training the initial model...” which further describes the concept is mere gathered data under prong 2 (insignificant extra solution activity— MPEP 2106.06g) and WURC under 2b (using gather data - MPEP 2106.05d). The limitations as recited in dependent claim 11 recites, “wherein the initial model is selected from …models...” which further describes the concept is mere gathered data under prong 2 (insignificant extra solution activity— MPEP 2106.06g) and WURC under 2b (using gather data - MPEP 2106.05d). The limitations as recited in dependent claim 9 recites, “…predicting whether the merged data satisfies the desired data quality level … generating new synthetic data …generating the merged dasta………” which further describes the concepts performed in the human mind including an observation, evaluation, judgment, and opinion, in step 2A prong one. Examiner’s Note 6. Synthetic data (According to Google): “Synthetic data is artificially generated information that mimics the mathematical and statistical patterns of real-world data, rather than being collected from actual observations. It is primarily created using algorithms and AI models (like Large Language Models or GANs) to train, test, and improve other software and systems “ Wang et al, US 20260087326, [Wang: Title and Abstract (“Generating synthetic data” and “eceiving data having missing elements having missing elements; evaluating the data for patterns with respect to the missing elements; labeling the data according to the patterns; generating labeled synthetic data using the labeled data; and inserting blanks into the labeled synthetic data according to associated labels of the labeled data to generate synthetic data with corresponding missing elements”)]. Siraswamy et al, US 20230086307, [Siraswamy: Abstract (“Data transformation and data quality checking is provided by reading data from a source datastore and storing the data into memory, performing in-memory processing of the data stored in memory, where the data is maintained in-memory for performance of the in-memory processing thereof, and where the in-memory processing includes performing one or more transformations on the data stored in memory, in which the data stored in memory is transformed and stored back into the memory and applying one or more data quality rules to the data stored in-memory, and based on performing the in-memory processing of the data stored and maintained in memory for the in-memory processing, loading to a target datastore at least some of the data processed by the in-memory processing”)] [Siraswamy: Paragraph 42 (“n addition, the data quality metadata repository 210 provides both a rules repository in which data quality rules are maintained and an information store for storing operational, audit, feedback, status, and historical data quality metrics information for a past runs. Another use of the information in the repository 210 is to provide data quality tracking and auditing for reports and/or application (software) 218. Further, aspects can use artificial intelligence (AI)/machine learning models to assist data stewards in creating new rules and modifying existing new rules, for instance recommending rules and attributes thereof, both before and/or during transformation/data quality check processing. A machine learning process can train the models from historical and other information, such as information in a data quality metadata repository. The training can train the models to identify data quality rules and attributes thereof to apply in varying situations with varying types of datasets and properties thereof. A trained ML model can be applied to catalog information about the source/target dataset(s) and optionally other information from the data quality metadata repository and generate configuration recommendations about data quality rule(s) to apply to a given source/target transformation/data quality check job. These configuration recommendations that then be provided to a user (such as a data steward) con configuring the data quality rule(s). These recommendations can be provided before, during, or after job runs. In particular embodiments, the ML model is applied during a job based on/after data quality failures, and the generated recommendations provide recommendations to a user as to how to change data quality rules(s) in response to the data quality failures”)] [Siraswamy: Paragraph 42 (“These configuration recommendations that then be provided to a user (such as a data steward) con configuring the data quality rule(s.”, i.e., ‘desired data quality level’)] [Siraswamy: Paragraph 42 (“for instance recommending rules and attributes thereof, both before and/or during transformation/data quality check processing. A machine learning process can train the models from historical and other information, such as information in a data quality metadata repository”, i.e., “the models from historical and other information, such as information in a data quality metadata repository” = ‘an initial model, which is pre-trained’)] [Siraswamy: Paragraph 42 (“A trained ML model can be applied to catalog information about the source/target dataset(s) and optionally other information from the data quality metadata repository and generate configuration recommendations about data quality rule(s) to apply to a given source/target transformation/data quality check job”, i.e., “A trained ML model” = ‘an initial model, which is pre-trained’)] [Siraswamy: Paragraph 43 (“Data stewards 302 define (304) rules in the data quality repository (e.g. 210 of FIG. 2) in which data quality rules are written and managed. The data stewards 302 can create rules, change rule parameters/thresholds, specify actions to take on rule failure/success, etc. Rules can be reusable/mapped for use for data quality checking of data from different datasets, potentially with different attributes being defined for any given rule depending on the dataset(s) involved. Additionally, the data stewards can set other configurations, for instance thresholds for how long data quality engines are to wait for steward/user corrective actions to rules and/or data on failure.”)]. Ezrielev et al, US 20250005183, [Ezrielev: Abstract (“To manage operation of a data pipeline when a portion of data is inaccessible may require generating synthetic portion of data to generalize the inaccessible portion of data. Prior to the generation of the synthetic portion of data, an analysis of an intended use of the inaccessible portion of data may be performed. The analysis may reduce the likelihood of generation and use of synthetic portion of data that is unreliable. Once obtained, the synthetic portion of data may be analyzed to determine a likelihood that the synthetic portion of the data may successfully generalize the inaccessible portion of data. When the synthetic portion of data is determined to meet or exceed quality criteria, the synthetic portion of data may be utilized by the data pipeline”)] [Ezrielev: Paragraph 18 (“The data pipeline may utilize one or more data processing systems to manage the operation of the data pipeline, which may include synthetic data generation when the data is determined to be inaccessible or limited to disclosure. Synthetic data generation may include using a trained inference model to generate synthetic data (e.g., prediction of the inaccessible data) for which may be implemented in the data pipeline when the requested data is inaccessible. For example, synthetic data generation may be implemented in order to provide synthetic data to a consumer when the requested data is inaccessible. By doing so, embodiments disclosed herein may provide a system for generating synthetic data when the data requested is inaccessible. The generation of synthetic data may reduce failures of the data pipeline (e.g., due to the inability to provide requested data) as a result of inaccessible data in the data pipeline”)]. Gokalp et al, US 11,868,436, [Gokalp: Column 1, lines 30-51 (“it is sometimes hard to determine the number of training examples that may eventually be required to train a classifier that meets targeted quality requirements, since the extent to which different examples assist in the model's learning may differ. As a result of these and other factors, generating a training set of labeled examples may often represent a non-trivial technical and resource usage challenge”)]. Han, US 20230177394, [Han: Abstract and paragraph 10 (“Provided is a method of generating training data of a manipulator robot and an apparatus performing the method. The method includes obtaining external appearance information of the plurality of components included in the training data from model information of the product and generating synthetic data in which position information of the component is labeled on a multi-angle image of the plurality of components, assembly information of the plurality of components, and assembly operation information of the plurality of components of the robot included in the training data based on the external appearance information”)]. Lee et al, US 20240232717, [Lee: Paragraphs 15 and 22 (“a machine-learning unit for training a machine-learning model in response to a request from the query-processing unit and generating synthetic spatiotemporal data based on the machine-learning model, and a data storage unit for storing raw spatiotemporal data and the generated synthetic spatiotemporal data, and the raw spatiotemporal data may be stored in the form of a table including an identifier column and a position column”)] [Lee: Paragraphs 29-30 (“determining whether the synthetic spatiotemporal data and the trained machine-learning model are present may comprise, when synthetic data corresponding to the target data and columns to be queried is not present, determining whether a machine-learning model corresponding to the target data and columns to be queried is present”)] [Lee: Paragraph 69 (“The machine-learning service provider 110 includes a machine-learning service module 111 for receiving a request for a machine-learning task and providing a result by communicating with the query-processing engine 100, a machine-learning execution module 112 for generating an ML model by training an ML model structure by accessing raw data in the data storage specified in the request for the task, and a synthetic data generation module 113 for generating synthetic data from the trained ML model. Also, there are model type storage 114 for storing ML model structures and model storage 115 for storing a model that is trained for specific data in the machine-learning execution module”)] [Lee: Paragraph 83 (“The query analysis module 102 checks whether a model available for generation of synthetic data is present by referring to the catalog storage 105 at step S401. When the model is not present, the process is terminated after an error is returned at step S402, whereas when the model is present, a request to generate synthetic data is made to the machine-learning service module 111 of the machine-learning service provider 110. In response to the request, the machine-learning execution module 112 loads the pretrained model and the model type thereof from the model storage 115 and the model type storage 114, respectively, at step S403 and generates synthetic data by executing a function that generates data from the machine-learning model at step S404. When the user's request includes time/space constraints, whether the generated data satisfies the constraints is checked at step S405. When the constraints are not satisfied, whether the generated data falls within a temporal/spatial range smaller than the temporal/spatial range specified in the constraints is checked at step S406, and data records only for the temporal/spatial range within which data is scarce are generated by setting conditions at steps S407 and S404”)]. Jandial et al, US 20240330682, [Jandial: Abstract and paragraph 3 (“Systems and methods for generating synthetic tabular data for machine learning and other applications are provided. In some embodiments, a variational autoencoder is trained to learn inter-feature correlations found in tabular data collected from real data sources. The trained variational autoencoder is used to train a generator model of a Generative Adversarial Network (GAN) to generate synthetic tabular data that exhibits the inter-feature correlation distribution found in the tabular data collected from real data sources. In some embodiments, processing devices perform operations comprising: receiving a set of tabular data records, each record comprising a plurality of features; training a first machine learning model using the tabular data records to learn correlations between the plurality of features; and training a second machine learning model, using the first machine learning model, to generate a synthetic tabular data records based at least on the one or more correlations between the plurality of feature”, i.e., “first machine learning model” = ‘an initial model’)] [Jandial: Paragraph 17 (“directed to machine learning related technologies for generating synthetic tabular data that closely replicates the inter-feature correlations found in tabular data collected from real data sources. Although real tabular data collected from real data sources is used to train machine learning models that generate synthetic tabular data, the resulting synthetic tabular data does not comprise records corresponding to any real individual and therefor may be used as training data to train other machine learning models without privacy risks”, i.e., “tabular data collected from real data sources” or “real tabular data collected from real data sources” = ‘original data’)]. Fortkort et al, US 20240395388, [Fortkort: Paragraph 186 (“The synthetic data is combined with the real data to create an augmented training dataset. This augmented dataset is used to train the AI that's responsible for generating PBM treatment recommendations. By training on a larger and more diverse dataset, the AI may be able to learn more robust and generalizable patterns, potentially improving its performance.”)]. Narayanan et al, US 20240249158, [Narayanan: Abstract (“machine learning monitoring and retraining techniques for automatically triggering model retraining based on evaluation scores. The techniques include receiving a request to process an input data object with a target machine learning model that is previously trained using an at least partially synthetic training dataset. The techniques include identifying a synthetic data object from the training dataset that corresponds to the input data object and, in response, modifying a holistic evaluation score for the model, initiating the performance of a labeling process for assigning a ground truth label to the input data object, and augmenting a supplemental training dataset with the input data object and the ground truth label. In the event that the holistic evaluation score decreased beyond a threshold, the model may be retrained with the supplemental training dataset”)] [Narayanan: Paragraph 77 (“the term “evaluation score” refers to a discrete component of a holistic evaluation data entity for a target machine learning model. An evaluation score may include any data type, format, and/or value that evaluates a particular aspect of the target machine learning model. By way of example, a first evaluation score may be dependent on a training dataset used to train the target machine learning model, a second evaluation score may be dependent on a performance of the target machine learning model, a third evaluation score may be dependent on one or more particular predictive outputs of the target machine learning model, and/or the like. The type, format, and/or value of an evaluation score may be based on the prediction domain. As an example, in an auto-adjudication prediction domain, an evaluation score may describe a biasness and/or fairness of a training dataset, a machine learning model, machine learning model decisions, and/or the like. An evaluation score, for example, may include a data evaluation score, a model evaluation score, a decision evaluation score, among others”)] [Narayanan: Paragraph 168 (“an evaluation score is a discrete component of a holistic evaluation data entity for the target machine learning model 302. An evaluation score may include any data type, format, and/or value that evaluates a particular aspect of the target machine learning model 302. Each evaluation score, by itself, may represent a quality of a different aspect of the target machine learning model 302 that, when taken together, may illustrate a holistic quality of the model. By way of example, a first evaluation score, such as the data evaluation score 306, may be dependent on the training dataset 304 used to train the target machine learning model 302 and may represent a quality of the training dataset 304. A second evaluation score, such as the model evaluation score 314, may be dependent on a performance of the target machine learning model 302 and may represent a quality of the configuration of the target machine learning model 302. A third evaluation score, such as the decision evaluation score 316, may be dependent on one or more particular predictive outputs of the target machine learning model and may represent a quality of predictive outputs generated by the target machine learning model 302 with respect to at least one predictive output class.”)] [Narayanan: Paragraph 11 (“training dataset comprising a plurality of synthetic data objects and a plurality of historical data objects”, i.e., “training dataset” = ‘merged data’)]. 7. Any inquiry concerning this communication or earlier communications from the examiner should be directed to [Hung D. Le], whose telephone number is [571-270-1404]. The examiner can normally be communicated on [Monday to Friday: 9:00 A.M. to 5:00 P.M.]. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Apu Mofiz can be reached on [571-272-4080]. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, contact [800-786-9199 (IN USA OR CANADA) or 571-272-1000]. Hung Le 07/17/2026 /HUNG D LE/Primary Examiner, Art Unit 2161
Read full office action

Prosecution Timeline

Aug 11, 2025
Application Filed
Jul 21, 2026
Non-Final Rejection mailed — §101 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705279
DISPLAY APPARATUS, BACKGROUND MUSIC PROVIDING METHOD THEREOF AND BACKGROUND MUSIC PROVIDING SYSTEM
2y 6m to grant Granted Aug 11, 2026
Patent 12694009
TRACKING EVALUATION OF WORKLOAD STABILITY THROUGH PERFORMANCE INDEXING
2y 1m to grant Granted Jul 28, 2026
Patent 12682286
GENERATING OPPORTUNITY PROFILE INSIGHTS
3y 1m to grant Granted Jul 14, 2026
Patent 12681999
Permissions-Aware Search with User Suggested Results and Document Verification
2y 2m to grant Granted Jul 14, 2026
Patent 12681934
Efficient Merging of Tabular Data with Post-Processing Compaction
2y 0m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
90%
Grant Probability
96%
With Interview (+6.2%)
2y 4m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1092 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month