Prosecution Insights
Last updated: August 30, 2026
Application No. 18/455,775

DATA-PRIVACY-PRESERVING SYNTHESIS OF REALISTIC SEMI-STRUCTURED TABULAR DATA

Non-Final OA §101§103
Filed
Aug 25, 2023
Examiner
LE, HUNG VAN
Art Unit
2145
Tech Center
2100 — Computer Architecture & Software
Assignee
SAP SE
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
23 currently pending
Career history
3
Total Applications
across all art units

Statute-Specific Performance

§101
31.4%
-8.6% vs TC avg
§103
68.6%
+28.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103
CTNF 18/455,775 CTNF 101938 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Information Disclosure Statement 06-52 The information disclosure statement (IDS) submitted on 2023/08/28. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 101 07-04-01 AIA 07-04 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 1-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (an abstract idea) without reciting significantly more. Regarding independent claims 1, 8, and 15 Step 1 -- whether the claim falls within any statutory category. See MPEP 2106.03 Claim 1 is drawn to a computer-implemented method claim; claim 8 is drawn to a non-transitory computer-readable storage medium claim; and claim 15 is drawn to a system claim. Therefore, each of these claims falls under one of the four categories of statutory subject matter: claim 1 falls within the process/method category, claim 8 falls within the manufacture category, and claim 15 falls within the machine/system/apparatus category. Accordingly, claims 1, 8, and 15 satisfy Step 1 under 35 U.S.C. § 101. Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. Regarding independent claim 1 , the claim is directed to a computer-implemented method for training one or more machine learning (ML) models, the method being executed by one or more processors, comprising receiving a real data table, providing a synthetic structured table, providing a sampled data table, transmitting a prompt to a large language model (LLM) system, receiving synthetic unstructured data from the LLM system, providing an aggregate synthetic table, and training an ML model using the aggregate synthetic table. The limitations of “providing a synthetic structured table based on the real data table,” “providing a sampled data table comprising a sub-set of real data of the real data table,” “transmitting a prompt to a large language model (LLM) system, the prompt being generated based on the real data table and the synthetic structured data table,” and “providing an aggregate synthetic table that includes at least a portion of the synthetic unstructured data” recite concepts that fall within the mental processes grouping of abstract ideas because, under their broadest reasonable interpretation, these limitations encompass observing, evaluating, selecting, organizing, and arranging information from data tables and generated data. These are concepts that can practically be performed in the human mind, including by observation, evaluation, judgment, and opinion. See MPEP § 2106.04(a)(2), subsection III . The limitation of “providing a synthetic structured table based on the real data table” also recites, or at least encompasses under the broadest reasonable interpretation in light of the specification, generating synthetic structured data based on statistics, distributions, or relationships determined from the real data table. Such generation of synthetic structured tabular data encompasses mathematical calculations and mathematical relationships, which fall within the mathematical concepts grouping of abstract ideas. See MPEP § 2106.04(a)(2), subsection I . The limitation of “training a ML model using the aggregate synthetic table” recites training a machine learning model using training data. Under the broadest reasonable interpretation in view of the specification and the ordinary meaning of training an ML model, the limitation encompasses using training data to determine or adjust model parameters through mathematical or statistical learning operations. Thus, this limitation falls within the mathematical concepts grouping of abstract ideas. See MPEP § 2106.04(a)(2), subsection I . The limitations of “receiving a real data table” and “receiving synthetic unstructured data from the LLM system” are data receipt limitations. These limitations are not relied upon here as the principal recited judicial exception at Step 2A Prong One, but may be considered as additional elements in the Step 2A Prong Two analysis. Accordingly, claim 1 recites abstract ideas including mental processes and mathematical concepts . Consistent with MPEP § 2106.04, subsection II , when a claim recites multiple abstract ideas, the limitations should be identified for the record and, where appropriate, considered together as a single abstract idea for further analysis. Therefore, claim 1 recites a judicial exception under Step 2A Prong 1. Independent claim 8 is drawn to a non-transitory computer-readable storage medium claim that recites operations substantially corresponding to the limitations of independent claim 1 , including receiving a real data table, providing synthetic and sampled data tables, transmitting a prompt to an LLM system, receiving synthetic unstructured data, providing an aggregate synthetic table, and training an ML model using the aggregate synthetic table. Accordingly, claim 8 recites the same abstract ideas identified for claim 1 , including mental processes and mathematical concepts , for similar reasons. See MPEP § 2106.04, subsection II and MPEP § 2106.04(a)(2), subsections I and III . Independent claim 15 is drawn to a system claim that recites a computing device and a computer-readable storage device having instructions that cause the computing device to perform operations substantially corresponding to the limitations of independent claim 1 , including receiving a real data table, providing synthetic and sampled data tables, transmitting a prompt to an LLM system, receiving synthetic unstructured data, providing an aggregate synthetic table, and training an ML model using the aggregate synthetic table. Accordingly, claim 15 recites the same abstract ideas identified for claim 1 , including mental processes and mathematical concepts , for similar reasons. See MPEP § 2106.04, subsection II and MPEP § 2106.04(a)(2), subsections I and III . Step 2A Prong 2 -- whether the claim as a whole integrates the recited judicial exception into a practical application of the exception or whether the claim is “directed to” the judicial exception. This evaluation is performed by (1) identifying whether there are any additional elements recited in the claim beyond the judicial exception, and (2) evaluating those additional elements individually and in combination to determine whether the claim as a whole integrates the exception into a practical application. See MPEP 2106.04(d). Regarding independent claim 1 , the claim recites additional elements of “one or more processors,” “receiving a real data table,” “transmitting a prompt to a large language model (LLM) system,” “receiving synthetic unstructured data from the LLM system,” and use of an LLM system and an ML model . The limitation of “receiving a real data table” is recited at a high level of generality and merely obtains input data for use in the recited data-generation and data-organization process. The limitation therefore amounts to mere data gathering, which is insignificant extra-solution activity. See MPEP § 2106.05(g) . The limitation of “receiving synthetic unstructured data from the LLM system” similarly recites receiving output data from the LLM system for use in later organizing or aggregating data. This limitation is recited at a high level of generality and amounts to insignificant extra-solution activity. See MPEP § 2106.05(g) . The limitations of “the method being executed by one or more processors,” “transmitting a prompt to a large language model (LLM) system,” and the recited use of the LLM system amount to no more than using generic computer or AI components as tools to perform the abstract idea. The claim does not recite a particular processor architecture, LLM architecture, inference mechanism, prompt-processing mechanism, or technical improvement to the LLM system. These limitations therefore amount to mere instructions to apply the judicial exception using generic computer components or generic AI tools. See MPEP § 2106.05(f) . The recited use of the ML model also does not integrate the judicial exception into a practical application. The claim recites “training a ML model using the aggregate synthetic table” without specifying a particular model architecture, loss function, optimization algorithm, training objective, parameter-update rule, deployment arrangement, or technical use of the trained model. This limitation merely links the abstract data-synthesis concept to the technological environment of ML training. See MPEP § 2106.05(h) . Claim 1 also does not recite an improvement to the functioning of a computer or to another technology or technical field. Although the claim is in the field of synthetic data generation and ML training, the claim does not recite a particular technical mechanism that improves processor operation, LLM operation, ML model operation, data storage, network communication, or another computer technology. The claim therefore does not integrate the judicial exception into a practical application under MPEP §§ 2106.04(d)(1) and 2106.05(a) . Accordingly, the additional elements, considered individually and in combination, do not integrate the recited judicial exception into a practical application. Rather, the claim recites data gathering and output receipt, generic computer implementation, and use of generic AI/ML tools to apply the abstract idea. Therefore, claim 1 is directed to the recited judicial exception under Step 2A Prong 2 . Regarding independent claim 8 , this claim is drawn to a non-transitory computer-readable storage medium claim reciting operations substantially corresponding to those recited in independent claim 1 . The overlapping operative limitations are rejected under the same rationale discussed above with respect to claim 1 . Claim 8 further recites the additional elements of “a non-transitory computer-readable storage medium coupled to one or more processors,” “instructions stored thereon,” and instructions that, “when executed by the one or more processors, cause the one or more processors to perform operations for training one or more machine learning (ML) models.” These additional elements do not integrate the recited judicial exception into a practical application. The non-transitory computer-readable storage medium, processors, and stored instructions are recited at a high level of generality and merely provide generic computer implementation for carrying out the abstract data-synthesis and ML-training operations. The claim does not recite a particular memory structure, processor architecture, instruction execution mechanism, or technical improvement to computer-readable media or processor operation. Therefore, these limitations amount to no more than mere instructions to apply the judicial exception using generic computer components. See MPEP § 2106.05(f) . These additional elements also merely link the abstract idea to the technological environment of a computer-readable storage medium and processor-executed instructions, without imposing a meaningful technological limitation on the abstract idea. See MPEP § 2106.05(h) . Accordingly, the additional elements of claim 8 , considered individually and in combination with the overlapping limitations analyzed for claim 1 , do not integrate the recited judicial exception into a practical application. Therefore, claim 8 is directed to the recited judicial exception under Step 2A Prong Two . See MPEP § 2106.04(d) . Regarding independent claim 15 , this claim is drawn to a system claim reciting operations substantially corresponding to those recited in independent claim 1 . The overlapping operative limitations are rejected under the same rationale discussed above with respect to claim 1 . Claim 15 further recites the additional elements of “a computing device,” and “a computer-readable storage device coupled to the computing device and having instructions stored thereon which, when executed by the computing device, cause the computing device to perform operations for training one or more machine learning (ML) models.” These additional elements do not integrate the recited judicial exception into a practical application. The computing device, computer-readable storage device, and stored instructions are recited at a high level of generality and merely provide generic computer components for implementing the abstract data-synthesis and ML-training operations. The claim does not recite a particular computing-device architecture, storage-device structure, processor configuration, memory arrangement, or technical improvement to computer operation. Therefore, these limitations amount to no more than mere instructions to apply the judicial exception using generic computer components. See MPEP § 2106.05(f) . These additional elements also merely link the abstract idea to the technological environment of a generic computer system, without imposing a meaningful technological limitation on the abstract idea. See MPEP § 2106.05(h) . Accordingly, the additional elements of claim 15 , considered individually and in combination with the overlapping limitations analyzed for claim 1 , do not integrate the recited judicial exception into a practical application. Therefore, claim 15 is directed to the recited judicial exception under Step 2A Prong Two . See MPEP § 2106.04(d) . Step 2B -- whether the claim amounts to significantly more than the judicial exception. See MPEP § 2106.05. Claims 1, 8, and 15 do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 1 , as discussed above, the additional elements include “one or more processors,” “receiving a real data table,” “transmitting a prompt to a large language model (LLM) system,” “receiving synthetic unstructured data from the LLM system,” and the recited use of an LLM system and ML model . The recited processor is a generic computer component used merely to apply the judicial exception. See MPEP § 2106.05(f) . The receiving and transmitting limitations are recited at a high level of generality and amount to mere data gathering or data transfer, which is insignificant extra-solution activity. See MPEP § 2106.05(g) . The LLM system and ML model are also recited generically and merely link the judicial exception to an AI/ML technological environment. See MPEP § 2106.05(h) . The receiving, transmitting, processor, storage, and instruction-execution functions are well-understood, routine, and conventional when recited at this level of generality. This finding is supported by the specification’s disclosure of generic computer implementation components, including processors, memory, storage devices, input/output devices, and networked computer systems. See specification 0074–0079; see also MPEP §§ 2106.05(d) and 2106.07(a)(III) . Regarding claim 8 , the claim recites substantially the same operative limitations as claim 1 and is rejected under the same rationale. Claim 8 additionally recites a non-transitory computer-readable storage medium , one or more processors , and stored instructions executed by the processors. These limitations are generic computer-implementation elements and do not add significantly more than the judicial exception. See MPEP §§ 2106.05(f), 2106.05(h) . Regarding claim 15 , the claim recites substantially the same operative limitations as claim 1 and is rejected under the same rationale. Claim 15 additionally recites a computing device , a computer-readable storage device , and stored instructions executed by the computing device. These limitations are generic computer-system elements and do not add significantly more than the judicial exception. See MPEP §§ 2106.05(f), 2106.05(h) . Even when considered individually and in combination, the additional elements of claims 1, 8, and 15 amount to generic computer implementation, insignificant extra-solution activity, and field-of-use or technological-environment limitations. Therefore, the claims do not recite an inventive concept and are not patent eligible under 35 U.S.C. § 101. Regarding dependent claims 2-7, 9-14, and 16-20 Step 1 -- whether the claim falls within any statutory category. See MPEP § 2106.03. Dependent claims 2–7 depend from independent claim 1 and are drawn to method claims . Therefore, claims 2–7 fall within the process/method category of statutory subject matter. Dependent claims 9–14 depend from independent claim 8 and are drawn to non-transitory computer-readable storage medium claims. Therefore, claims 9–14 fall within the manufacture category of statutory subject matter. Dependent claims 16–20 depend from independent claim 15 and are drawn to system claims . Therefore, claims 16–20 fall within the machine/product/apparatus category of statutory subject matter. Accordingly, dependent claims 2–7, 9–14, and 16–20 each fall under one of the four categories of statutory subject matter: process/method, machine/product/apparatus, manufacture, or composition of matter . Step 2A Prong 1 – whether the claim recites a judicial exception. See MPEP 2106.04, subsection II. Regarding dependent claims 2, 9, and 16 , these claims recite the limitation of “wherein the prompt comprises rows of the sampled data table as few-shot examples for a LLM of the LLM system to generate the synthetic unstructured data.” This limitation recites selecting and arranging rows of sampled data as examples in a prompt. Under its broadest reasonable interpretation, this limitation encompasses observation, evaluation, judgment, and selection of information, which can practically be performed in the human mind. Thus, this limitation recites a mental process under MPEP § 2106.04(a)(2), subsection III . Regarding dependent claims 3, 10, and 17 , these claims recite the limitation of “wherein the prompt is generated using a prompt template.” This limitation recites organizing information according to a predefined template. Under its broadest reasonable interpretation, this limitation can practically be performed in the human mind or with pen and paper by inserting selected information into a template. Thus, this limitation recites a mental process under MPEP § 2106.04(a)(2), subsection III . Regarding dependent claims 4, 11, and 18 , these claims recite the limitation of “wherein the sampled data table is provided by sampling rows of the real data table.” This limitation recites selecting rows from a real data table. Under its broadest reasonable interpretation, selecting rows from a table can practically be performed in the human mind by observation, evaluation, and judgment. Thus, this limitation recites a mental process under MPEP § 2106.04(a)(2), subsection III . Regarding dependent claims 5, 12, and 19 , these claims recite the limitation of “wherein providing an aggregate synthetic table comprises selectively filtering at least a portion of a semi-structured synthetic table that is provided from the LLM system.” This limitation recites evaluating and selectively retaining or removing portions of data from a semi-structured synthetic table. Under its broadest reasonable interpretation, such filtering can practically be performed in the human mind by observation, evaluation, and judgment. Thus, this limitation recites a mental process under MPEP § 2106.04(a)(2), subsection III . Regarding dependent claims 6, 13, and 20 , these claims recite the limitation of “wherein providing an aggregate synthetic table comprises aggregating at least portions of multiple semi-structured synthetic table.” This limitation recites combining and arranging portions of multiple semi-structured synthetic tables. Under its broadest reasonable interpretation, aggregating portions of tables can practically be performed in the human mind or with pen and paper. Thus, this limitation recites a mental process under MPEP § 2106.04(a)(2), subsection III . Regarding dependent claims 7 and 14 , these claims recite the limitation of “wherein the synthetic structured table comprises synthetic structured data that is generated based on one or more distributions determined from the real data table.” This limitation recites generating data based on one or more distributions determined from the real data table. A distribution is a mathematical/statistical relationship, and determining distributions from data and generating data based on those distributions recites mathematical relationships and mathematical calculations. Thus, this limitation recites a mathematical concept under MPEP § 2106.04(a)(2), subsection I . Accordingly, dependent claims 2–7, 9–14, and 16–20 recite additional limitations that fall within the mental processes and/or mathematical concepts groupings of abstract ideas. These dependent claims also incorporate the abstract ideas recited in their respective independent claims. Therefore, dependent claims 2–7, 9–14, and 16–20 recite judicial exceptions under Step 2A Prong One . Step 2A Prong 2 -- whether the claim as a whole integrates the recited judicial exception into a practical application of the exception or whether the claim is “directed to” the judicial exception. This evaluation is performed by (1) identifying whether there are any additional elements recited in the claim beyond the judicial exception, and (2) evaluating those additional elements individually and in combination to determine whether the claim as a whole integrates the exception into a practical application. See MPEP 2106.04(d). Regarding dependent claims 2, 9, and 16 , these claims recite the additional limitation of “wherein the prompt comprises rows of the sampled data table as few-shot examples for a LLM of the LLM system to generate the synthetic unstructured data.” This limitation merely specifies that rows of sampled data are used as examples in a prompt. The claim does not recite a particular prompt encoding mechanism, LLM architecture, inference-control process, or technical improvement to the LLM system. Thus, this limitation amounts to using the LLM as a tool to apply the judicial exception and generally linking the exception to an LLM technological environment. See MPEP §§ 2106.05(f) and 2106.05(h) . Regarding dependent claims 3, 10, and 17 , these claims recite the additional limitation of “wherein the prompt is generated using a prompt template.” This limitation merely specifies organizing information according to a template. The claim does not recite a particular computer-specific template structure, prompt compiler, memory arrangement, or technical improvement to computer or LLM functionality. Thus, this limitation amounts to organizing information and using generic computer/LLM implementation to apply the judicial exception. See MPEP §§ 2106.05(f) and 2106.05(h) . Regarding dependent claims 4, 11, and 18 , these claims recite the additional limitation of “wherein the sampled data table is provided by sampling rows of the real data table.” This limitation merely specifies selecting rows from a real data table. The claim does not recite a particular sampling algorithm, database-access technique, data structure, or technical improvement to computer operation. Thus, this limitation amounts to data selection or data gathering for use in the abstract idea. See MPEP §§ 2106.05(f), 2106.05(g), and 2106.05(h) . Regarding dependent claims 5, 12, and 19 , these claims recite the additional limitation of “wherein providing an aggregate synthetic table comprises selectively filtering at least a portion of a semi-structured synthetic table that is provided from the LLM system.” This limitation recites filtering data at a high level of generality. The claim does not recite a particular filtering algorithm, similarity metric, privacy threshold, vector-comparison process, or technical mechanism for preventing private-data leakage. Thus, the limitation recites the result of filtering data without a particular technical way of achieving that result and does not integrate the judicial exception into a practical application. See MPEP §§ 2106.05(e), 2106.05(f), and 2106.05(h) . Regarding dependent claims 6, 13, and 20 , these claims recite the additional limitation of “wherein providing an aggregate synthetic table comprises aggregating at least portions of multiple semi-structured synthetic table.” This limitation merely specifies combining portions of multiple synthetic tables. The claim does not recite a particular database architecture, table-join mechanism, indexing technique, memory layout, or technical improvement to computer storage or processing. Thus, this limitation amounts to organizing information using generic computer implementation. See MPEP §§ 2106.05(f) and 2106.05(h) . Regarding dependent claims 7 and 14 , these claims recite the additional limitation of “wherein the synthetic structured table comprises synthetic structured data that is generated based on one or more distributions determined from the real data table.” This limitation recites generating synthetic data based on mathematical/statistical distributions. The limitation is part of the mathematical concept itself and does not add an additional element that integrates the judicial exception into a practical application. See MPEP §§ 2106.04(d), 2106.05(f), and 2106.05(h) . Accordingly, dependent claims 2–7, 9–14, and 16–20 , considered individually and as a whole with their respective independent claims, do not integrate the recited judicial exceptions into a practical application. The dependent limitations merely narrow the abstract data-selection, prompt-generation, filtering, aggregation, and statistical data-generation concepts, or generally apply those concepts in a generic LLM/ML computer environment. Therefore, dependent claims 2–7, 9–14, and 16–20 are directed to the recited judicial exceptions under Step 2A Prong 2 . Step 2B -- whether the claim amounts to significantly more than the judicial exception. See MPEP § 2106.05. Dependent claims 2–7, 9–14, and 16–20 do not include additional elements that are sufficient to amount to significantly more than the judicial exception. These claims incorporate the limitations of their respective independent claims and add limitations that further narrow the same abstract data-selection, prompt-generation, filtering, aggregation, statistical data-generation, and ML-training concepts. Regarding claims 2, 9, and 16 , the limitation of “wherein the prompt comprises rows of the sampled data table as few-shot examples for a LLM of the LLM system to generate the synthetic unstructured data” merely specifies using selected rows as examples in a prompt. The limitation does not recite a particular LLM architecture, tokenization mechanism, prompt encoding mechanism, inference-control technique, or improvement to computer or LLM functionality. Thus, the limitation does not add significantly more than the judicial exception. See MPEP §§ 2106.05(f) and 2106.05(h) . Regarding claims 3, 10, and 17 , the limitation of “wherein the prompt is generated using a prompt template” merely specifies organizing information according to a template. The limitation does not recite a computer-specific template structure, specialized prompt compiler, memory arrangement, or technical improvement to LLM or computer functionality. Thus, the limitation does not add significantly more than the judicial exception. See MPEP §§ 2106.05(f) and 2106.05(h) . Regarding claims 4, 11, and 18 , the limitation of “wherein the sampled data table is provided by sampling rows of the real data table” merely specifies selecting rows from a real data table. The limitation does not recite a particular sampling algorithm, database-access technique, indexing structure, or technical improvement to computer operation. To the extent this limitation is considered separately from the abstract idea, it is insignificant extra-solution activity. See MPEP § 2106.05(g) . Regarding claims 5, 12, and 19 , the limitation of “wherein providing an aggregate synthetic table comprises selectively filtering at least a portion of a semi-structured synthetic table that is provided from the LLM system” merely recites filtering data at a high level of generality. The limitation does not recite a particular filtering algorithm, similarity metric, privacy threshold, vector-comparison process, or technical mechanism for preventing private-data leakage. Therefore, the limitation recites the result of filtering data without a particular technical way of achieving that result and does not add significantly more than the judicial exception. See MPEP §§ 2106.05(e), 2106.05(f), and 2106.05(h) . Regarding claims 6, 13, and 20 , the limitation of “wherein providing an aggregate synthetic table comprises aggregating at least portions of multiple semi-structured synthetic table” merely specifies combining portions of multiple synthetic tables. The limitation does not recite a particular database architecture, table-join mechanism, indexing technique, memory layout, or technical improvement to computer storage or processing. Thus, the limitation does not add significantly more than the judicial exception. See MPEP §§ 2106.05(f) and 2106.05(h) . Regarding claims 7 and 14 , the limitation of “wherein the synthetic structured table comprises synthetic structured data that is generated based on one or more distributions determined from the real data table” further narrows the mathematical concept by specifying generation of synthetic structured data based on distributions determined from the real data table. The limitation does not recite a particular technical data-generation mechanism, privacy-preserving statistical protocol, memory structure, processor configuration, or improvement to computer operation. Thus, the limitation does not add significantly more than the judicial exception. See MPEP §§ 2106.05 and 2106.05(f)–(h) . The generic computer-related elements incorporated through claims 1, 8, and 15 , including processors, computer-readable storage media, computing devices, stored instructions, receiving data, transmitting data, LLM usage, and ML-model usage, are well-understood, routine, and conventional when recited at this level of generality. This finding is supported by the specification’s disclosure of generic processors, memory, storage devices, input/output devices, APIs, LLM systems, ML systems, and networked computer systems. See MPEP §§ 2106.05(d) and 2106.07(a)(III) . Even when considered individually and in combination with their respective parent claims, dependent claims 2–7, 9–14, and 16–20 do not add an inventive concept. The dependent limitations merely narrow the judicial exceptions or apply them using generic computer/AI components. Accordingly, claims 2–7, 9–14, and 16–20 do not amount to significantly more than the judicial exception and are not patent eligible under 35 U.S.C. § 101. Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-103 AIA The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. 07-23-aia AIA The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 07-21-aia AIA Claim 1-4, 7-11, a nd 14-18 are reject ed under 35 U.S.C. §103 as being unpatentable over Bruera et al. ( Bruera ), Non-Patent Literature, “Generating Realistic Synthetic Curricula Vitae for Machine Learning Applications under Differential Privacy”, published on June 2022 and cited in the IDS filed on 8/28/23, and relied upon at pages 53-63, in view of Jandial et al. (Jandial) , U.S. Patent Application Publication No. US2024/0330682A1, and further in view of Hegselmann et al. (Hegselmann) , Non-Patent Literature, “TabLLM: Few-shot Classification of Tabular Data with Large Language Models”, published on April 2023, and relied upon at pages 1-4 and Supplementary Materials page 1. As to in dependent claim 1 , Bruera teaches a computer-implemented method for training one or more machine learning models, the method being executed by one or more processors, comprising: “transmitting a prompt to a large language model (LLM) system, the prompt being generated based on the real data table and the synthetic structured data table;” “receiving synthetic unstructured data from the LLM system;” “providing an aggregate synthetic table that includes at least a portion of the synthetic unstructured data;” and “training a ML model using the aggregate synthetic table.” Bruera teaches generating realistic synthetic curricula vitae (“CVs”) for machine learning applications under differential privacy. Bruera explains that applications involving machine learning in Human Resources require training data, but CVs contain sensitive personal information, and Bruera therefore proposes generating synthetic CVs that preserve useful real-world distributions while providing privacy guarantees ( Bruera , p. 53, Abstract; p. 54, left column, Introduction). Bruera further teaches that the generated synthetic CVs can be used for training machine learning models instead of, or together with, the original data ( Bruera , p. 53, Abstract; p. 54, left column, Introduction). With respect to “transmitting a prompt to a large language model (LLM) system, the prompt being generated based on the real data table and the synthetic structured data table,” Bruera teaches guiding text generation of a Transformer-based generative language model with a manually-prepared set of prompts into which attributes sampled from a Bayesian network are plugged ( Bruera , p. 53, Abstract). Bruera further teaches that candidate attributes sampled from the Bayesian network are used to generate text for each section of a synthetic CV, and that the attributes are plugged into linguistic structures called prompts, which are then provided as inputs to the generative model ( Bruera , p. 56, left column, §3.3). Bruera also teaches that the generation uses GPT-2, where generation starts from a prompt for section 1, generated text is appended, and the next prompt is fed back to GPT-2 until all CV sections are generated ( Bruera , p. 58, left column, §4.4, para. beginning “As per GPT-2…”). Thus, Bruera teaches transmitting prompts to a generative language model and generating those prompts based on structured candidate-attribute data. With respect to “receiving synthetic unstructured data from the LLM system,” Bruera teaches receiving generated natural-language text from the generative language model. Bruera explains that the prompts are provided as starting points for generation of each section of the CV and that the generative model generates natural-language sections of the synthetic CV ( Bruera , p. 56, left column, §3.3). Bruera further explains that GPT-2 generation process works cyclically by generating text from the prompt for section 1, appending the generated result to the prompt for section 2, feeding the combined text back as a new prompt to GPT-2, and continuing “until the end of all the sections is reached, and the CV is ready” ( Bruera , p. 58, left column, para. 1). With respect to “providing an aggregate synthetic table that includes at least a portion of the synthetic unstructured data,” Bruera teaches generating a set of synthetic CVs using the disclosed process. Specifically, Bruera states that the authors “generate in this way a set of 4000 synthetic CVs” ( Bruera , p. 58, left column, para. 1). The generated set of synthetic CVs includes structured candidate attributes and generated natural-language text sections. Under the broadest reasonable interpretation, such a set of structured records including generated text corresponds to an aggregate synthetic dataset/table including at least a portion of synthetic unstructured data. With respect to “training a ML model using the aggregate synthetic table,” Bruera teaches evaluating whether synthetic CVs can be used for downstream machine learning in HR applications. Bruera expressly states that extrinsic evaluation is performed by training a model for a classification task using the synthetic CVs ( Bruera , p. 53, Abstract). Bruera further teaches using a Candidate Role Classification task, splitting real CVs into train and test data, and training/testing two classifiers, fastText and BERT, using different training data, including generated synthetic CVs and augmented training data ( Bruera , p. 59, §5.2.1–5.2.2). Bruera teaches extracting candidate attributes from real CV samples and transforming the CVs into structured data entries (Bruera, p. 55, left column, §3, para. 1; p. 55, left column, §3.1, para. 1), learning conditional dependencies and conditional probability distributions using a Bayesian network, and sampling synthetic candidate attributes from that Bayesian network ( Bruera , p. 55, left column, §3.2, para. 1; p. 55, right column, §3.2, para. 1), and sampling synthetic candidate attributes from the Bayesian network for synthetic CV generation ( Bruera , p. 55, right column, §3.2, last para. continuing to p. 56, left column, §3.3, para. 1). However, Bruera does not expressly teach “receiving a real data table” and “providing a synthetic structured table based on the real data table” in the direct tabular-data form recited by claim 1. In the same field of endeavor, Jandial teaches “receiving a real data table” and “providing a synthetic structured table based on the real data table.” Jandial is directed to systems and methods for generating synthetic tabular data for machine learning and other applications ( Jandial , front page, title). Jandial teaches receiving a set of tabular data records, where each tabular data record comprises a plurality of features ( Jandial , front page Abstract; Fig. 4, step 410). Jandial further teaches training a first machine learning model using the set of tabular data records to learn one or more correlations between the plurality of features ( Jandial , Fig. 4, step 412). Jandial further teaches training a second machine learning model, using the first machine learning model, to generate a set of synthetic tabular data records based at least on one or more correlations between the plurality of features ( Jandial , Fig. 4, step 414). Jandial also teaches executing a machine learning model to generate a set of synthetic tabular data records and storing the set of synthetic tabular data records ( Jandial , Fig. 8, steps 810–814). Thus, Jandial teaches receiving a real data table and providing a synthetic structured table based on the real data table. Bruera and Jandial are analogous to the claimed invention because both are in the same field of endeavor of generating synthetic data from real data for use in machine learning applications, including privacy-preserving or privacy-risk-reducing use of synthetic data. Bruera teaches generating privacy-preserving synthetic CVs for machine learning applications using structured attributes and generated text, while Jandial teaches generating synthetic tabular data from real tabular data by learning correlations among tabular features. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine Bruera ’s prompt-guided synthetic CV generation method with Jandial ’s synthetic tabular-data generation techniques. The motivation to combine is provided by Jandial , which recognizes that collecting real tabular data may raise privacy concerns because real data may include personal information traceable to real persons, and that synthetic tabular data can be used to train machine learning models while avoiding privacy risks and maintaining inter-feature correlations useful for machine learning ( Jandial , spec. pp. 1–2, 0001–0020). A person of ordinary skill in the art would have been motivated to incorporate Jandial ’s tabular-data correlation-learning and synthetic structured-data generation into Bruera ’s synthetic-data generation pipeline to improve the statistical fidelity of the structured attributes used to guide generation of synthetic text while preserving privacy. The combination of Bruera and Jandial , however, does not expressly teach “providing a sampled data table comprising a sub-set of real data of the real data table.” In the same field of endeavor, Hegselmann teaches “providing a sampled data table comprising a sub-set of real data of the real data table.” Hegselmann teaches applying large language models to zero-shot and few-shot classification of tabular data by prompting the language model with a serialization of tabular data into a natural-language string together with a short description of the classification problem ( Hegselmann , p. 1, Abstract). Hegselmann further teaches that TabLLM uses “tabular data with k labeled rows,” serializes feature names and values into a natural-language string, adds a task-specific prompt, and fine-tunes the LLM using labeled examples ( Hegselmann , p. 2, Fig. 1). Hegselmann further states that a tabular dataset has rows and columns/features and that, for k-shot classification experiments, “we only use a subset Dk of size k—sampled from D with replacement—for fine-tuning or training” ( Hegselmann , p. 3, §3.1). Accordingly, Hegselmann teaches selecting a sampled subset of rows from a tabular dataset and using those sampled rows in connection with an LLM input. Bruera , Jandial , and Hegselmann are analogous to the claimed invention because all three references concern machine learning using structured, tabular, or semi-structured data, and Bruera and Hegselmann both involve language-model prompting based on structured or tabular information. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to further combine Hegselmann ’s sampled-row tabular prompting technique with the combined teachings of Bruera and Jandial . The motivation to combine is provided by Hegselmann , which teaches that tabular rows can be serialized into natural-language strings and combined with task-specific prompts for use by large language models in few-shot settings, using a small number of sampled labeled examples ( Hegselmann , pp. 1–3, Abstract, Fig. 1, §3.1). A person of ordinary skill in the art would have been motivated to use Hegselmann ’s sampled tabular-row technique in the Bruera / Jandial system to provide representative real tabular examples to the LLM, thereby improving prompt construction and helping the LLM generate synthetic unstructured text consistent with real tabular-data patterns. As to dependent Claim 2 , the combination of Bruera , Jandial , and Hegselmann teaches the method of Claim 1 as set forth above. Claim 2 further recites: “ wherein the prompt comprises rows of the sampled data table as few-shot examples for a LLM of the LLM system to generate the synthetic unstructured data.” Bruera teaches the remaining portion of the limitation requiring the LLM system “to generate the synthetic unstructured data.” Bruera teaches that “candidate attributes sampled from the Bayesian network together with the artificial personal details are then used to generate the text for each section of the synthetic CV” ( Bruera , p. 56, §3.3). Bruera further teaches that the attributes are plugged into “linguistic structures called prompts,” and that “[t]hese prompts, with the attributes plugged in, are what will be actually presented, one after another, to the generative model as starting points for the generation of each section of the CV” ( Bruera , p. 56, §3.3). Bruera also teaches using GPT-2, where “generation starts from the prompt for section 1,” the generated text is appended, and the resulting text is fed back as a new prompt to GPT-2 until “the CV is ready” ( Bruera , p. 58, left column, para. 1). Therefore, Bruera teaches using the LLM/generative language model to generate synthetic unstructured data. Bruera teaches using prompts for a generative language model to generate synthetic unstructured data, but Bruera does not expressly teach that the prompt comprises rows of a sampled data table as few-shot examples. Hegselmann further teaches wherein the prompt comprises rows of the sampled data table as few-shot examples for a LLM. Specifically, Hegselmann teaches applying large language models to few-shot classification of tabular data by prompting the LLM with serialized tabular data, stating: “We study the application of large language models to zero-shot and few-shot classification of tabular data” and “We prompt the large language model with a serialization of the tabular data to a natural-language string” ( Hegselmann , p. 1, Abstract). Hegselmann further teaches that “[i]n the few-shot setting, we fine-tune the large language model using some labeled examples” ( Hegselmann , p. 1, Abstract). Hegselmann also illustrates that the LLM prompt includes sampled table rows as examples. Figure 1 shows “Tabular data with k labeled rows,” followed by “Serialize feature names and values into natural-language string,” “Add task-specific prompt,” and “Fine-tune LLM using labeled examples” ( Hegselmann , p. 2, Fig. 1). Figure 1 further shows exemplary tabular rows with columns including “age,” “education,” “gain,” and “income,” and shows these row values converted into natural-language prompt text, e.g., “The age is 29. The education is Doctorate. The gain is 1086,” together with the prompt “Does this person earn more than 50000 dollars? Yes or no? Answer:” ( Hegselmann , p. 2, Fig. 1). Hegselmann further expressly teaches that the rows used in the few-shot setting are sampled from a tabular dataset. Hegselmann states: “Suppose we have a tabular dataset with n rows and d columns or features” and, “[f]or our k-shot classification experiments, we only use a subset Dk of size k—sampled from D with replacement—for fine-tuning or training” ( Hegselmann , p. 3, §3.1). Hegselmann also teaches that “[t]o use an LLM for tabular data, the table must be transformed into a natural text representation,” and that “when prompting an LLM, there is a template used to both serialize the inputs into one natural-language string, and to provide the prompt itself” ( Hegselmann , p. 3, §3.1). Thus, Hegselmann teaches a prompt comprising rows of a sampled data table as few-shot examples for an LLM. Bruera and Hegselmann are analogous to the claimed invention as both are from the same field of endeavor of using language models with structured, tabular, or semi-structured data in machine-learning applications. Bruera uses structured candidate attributes to prompt a generative language model to generate synthetic CV text, while Hegselmann teaches serializing sampled tabular rows into prompt form for LLM processing in few-shot settings. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the prompt-generation method of Bruera to include rows of a sampled data table as few-shot examples, as taught by Hegselmann , for the purpose of providing representative tabular examples to the LLM and improving the LLM’s ability to process prompt inputs based on tabular data. The motivation to combine is provided by Hegselmann , which teaches that LLMs can be prompted with serialized tabular data and can use a small number of labeled examples in a few-shot setting ( Hegselmann , pp. 1–3, Abstract, Fig. 1, §3.1), thereby providing a known technique for adapting LLM processing to tabular data examples. As to dependent Claim 3 , the combination of Bruera , Jandial , and Hegselmann teaches the method of Claim 1 as set forth above. Claim 3 further recites: “wherein the prompt is generated using a prompt template.” Bruera teaches this limitation. Specifically, Bruera teaches generating synthetic CVs “by guiding the text generation of a Transformer-based generative language model with a manually-prepared set of prompts where the attributes sampled from the Bayesian network are plugged in” ( Bruera , p. 53, Abstract). Bruera further teaches that candidate attributes sampled from the Bayesian network are used to generate text for each section of the synthetic CV, and that “we further plug them in a set of linguistic structures called prompts” ( Bruera , p. 56, §3.3). Bruera explains that the prompts are “typical bits of sentences where the attribute would be found in a human language,” such as “I studied x at y” and “I worked as p at q for r,” and that “[t]hese prompts, with the attributes plugged in, are what will be actually presented, one after another, to the generative model as starting points for the generation of each section of the CV” ( Bruera , p. 56, §3.3). Thus, Bruera teaches generating a prompt by inserting structured attribute values into predefined linguistic prompt forms. Bruera further expressly teaches using a template structure for prompt generation. In particular, Bruera teaches that, “[t]o generate the synthetic CVs with GPT-2, we manually define a template CV structure reflecting standard versions of CVs,” including sections for introductory fake personal information, education, work experience, linguistic skills, and hobbies ( Bruera , p. 57, §4.4). Bruera also teaches that “[f]or each section, we write a set of two to five possible prompts to randomly sample from at generation time,” and gives example prompt templates for education, including “I studied x at y” and “I attended y, where I studied x” ( Bruera , p. 57, §4.4). These disclosures teach that the prompt is generated using a prompt template. As to dependent Claim 4 , the combination of Bruera , Jandial , and Hegselmann teaches the method of Claim 1 as set forth above. Claim 4 further recites: “wherein the sampled data table is provided by sampling rows of the real data table.” Bruera teaches real source data that is transformed into structured data entries for subsequent synthetic-data generation. Specifically, Bruera teaches that its approach begins with “the extraction of candidate attributes from a dataset of real CVs, in fact transforming the CVs into structured data entries” ( Bruera , p. 55, §3). Bruera further teaches that “[t]he first step consists in the extraction of candidate attributes from a dataset of real CVs,” and that the output of this step is “a structured dataset, containing the key attributes which constitute a candidate profile” ( Bruera , p. 55, §3.1). Bruera also teaches that, “As a starting point, we use the dataset of real-world CVs,” and that “[t]he dataset, made of around 28000 CVs in English, was collected online from a dedicated website” ( Bruera , p. 56, §4.1). Thus, Bruera teaches real data that is converted into structured data entries. Jandial teaches the real data table as tabular records having rows and features. Specifically, Jandial teaches “receiving a set of tabular data records, each tabular data record of the set of tabular data records comprising a plurality of features” ( Jandial , Fig. 4, step 410; front page, Abstract). Jandial further teaches that collected tabular data records may be used to train a machine-learning model and may be insufficient in number, incomplete, noisy, or under-balanced, thereby indicating the use of collected real tabular records as the source data ( Jandial , p. 3, 0025). Jandial also teaches that “collected tabular data 205, in some embodiments, comprises a tabular data format wherein each row of the collected tabular data 205 define[s] a single tabular data record 260 and each of the columns 270 corresponds to an individual feature” ( Jandial , p. 4, 0031). Thus, Jandial teaches a real data table having rows. Hegselmann teaches providing a sampled data table by sampling rows from the real data table. Specifically, Hegselmann teaches “Tabular data with k labeled rows” ( Hegselmann , p. 2, Fig. 1). Hegselmann further teaches: “Suppose we have a tabular dataset with n rows and d columns or features,” and “[f]or our k-shot classification experiments, we only use a subset Dk of size k—sampled from D with replacement—for fine-tuning or training” ( Hegselmann , p. 3, §3.1). Hegselmann further confirms in the supplementary materials that “[e]ach dataset was separated into 80/20 train-test splits” and that “[t]he k labeled examples Dk were sampled in a class-balanced manner from the training set” ( Hegselmann , Supplementary Materials, p. 1, §1.1). Thus, Hegselmann expressly teaches sampling rows/examples from a tabular dataset to form a sampled subset. Bruera , Jandial , and Hegselmann are analogous to the claimed invention because each is directed to machine-learning processing of structured, tabular, or semi-structured data. Bruera teaches generating synthetic CV data from real CV data by extracting structured attributes and using those attributes in a generative language-model pipeline. Jandial teaches receiving real tabular data records and generating synthetic tabular data records based on learned correlations among features. Hegselmann teaches sampling rows from a tabular dataset and using those sampled rows in a few-shot LLM setting. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to use Hegselmann ’s row-sampling technique with the combined system of Bruera and Jandial so that representative rows from the real tabular dataset are selected as the sampled data table. The motivation to combine is provided by Hegselmann , which teaches that sampled subsets of tabular rows are used for k-shot LLM training or fine-tuning, thereby providing a known technique for selecting representative tabular examples for LLM-based processing ( Hegselmann , p. 3, §3.1; Supplementary Materials, p. 1, §1.1). As to dependent Claim 7 , the combination of Bruera , Jandial , and Hegselmann teaches the method of Claim 1 as set forth above. Claim 7 further recites: “wherein the synthetic structured table comprises synthetic structured data that is generated based on one or more distributions determined from the real data table.” Bruera teaches generating synthetic structured data based on distributions learned from real data. Specifically, Bruera teaches that its approach starts “from real samples, from which relevant attributes are extracted, and whose conditional dependencies and distributions across candidates are modelled through a Bayesian network (BN)” ( Bruera , p. 54, §1, left column, paragraph beginning “In our approach, we start from real samples…”). Bruera further teaches that “the structure and the conditional probability distributions are learnt from real samples,” and that the system generates “synthetic candidates in the form of sets of attributes, whose conditional distributions are close enough to those of real world candidates, yet providing DP” ( Bruera , p. 54, §1, left column, same paragraph). Bruera further teaches that the real CV data is transformed into structured data and used to build the distribution model. Specifically, Bruera states that the approach includes “the extraction of candidate attributes from a dataset of real CVs, in fact transforming the CVs into structured data entries,” followed by “the creation of a differentially private Bayesian network representing the conditional dependencies between the selected attributes,” and then “sampling a synthetic set of attributes for a candidate from the differentially private Bayesian network” ( Bruera , p. 55, §3, left column, paragraph beginning “Our approach is composed of three steps…”). Bruera also teaches that “[f]rom this structured dataset of candidate attributes, we build a Bayesian network, a probabilistic graphical model which represents a set of variables and their conditional dependencies” ( Bruera , p. 55, §3.2, left column, paragraph beginning “From this structured dataset…”). Bruera explains that each node in the Bayesian network outputs “the probability distribution on the node’s values,” and that “[t]his function constitutes a conditional probability distribution” ( Bruera , p. 55, §3.2, right column, paragraph beginning “Each node is associated…”). Bruera further teaches that “the conditional probability distributions are usually learnt from data” ( Bruera , p. 55, §3.2, right column, same paragraph). Bruera further teaches that the generated synthetic structured attributes preserve the statistical properties of the original dataset. Specifically, Bruera teaches that “generated values follow the conditional dependencies of the attributes and preserve the consistency and statistical properties of the original dataset up to the noise addition,” and that “we can generate a synthetic set of attributes for a realistic, but not real, candidate” ( Bruera , p. 56, continuation of §3.2, left column, paragraph beginning “generated values follow…”). Bruera also provides implementation-level support, stating that “the conditional dependencies among the nodes … are learnt under differential privacy using PrivBayes,” and that “[t]he conditional probabilities are learnt from the reduced set of around 1500 CVs” ( Bruera , p. 57, §4.3, right column, paragraph beginning “For the Bayesian network…”). Thus, Bruera teaches synthetic structured data generated based on one or more distributions determined from real structured data derived from the real CV dataset. Jandial further teaches synthetic structured/tabular data generated based on correlations or distributions determined from real tabular data. Specifically, Jandial teaches that “a variational autoencoder is trained to learn inter-feature correlations found in tabular data collected from real data sources,” and that the trained model is used to generate “synthetic tabular data that exhibits the inter-feature correlation distribution found in the tabular data collected from real data sources” ( Jandial , front page, Abstract). Jandial further teaches “receiving a set of tabular data records,” “training a first machine learning model using the set of tabular data records to learn one or more correlations between the plurality of features,” and “training a second machine learning model … to generate a set of synthetic tabular data records based at least on the one or more correlations between the plurality of features” ( Jandial , Fig. 4, steps 410–414). Jandial also teaches executing a machine-learning model to generate synthetic tabular data records, generating records comprising correlations between features, and storing the synthetic tabular records ( Jandial , Fig. 8, steps 810–814). Jandial provides additional paragraph-level support that the learned correlations are determined from real tabular data and preserved in the synthetic tabular records. Jandial teaches generating synthetic tabular data that “closely replicates the inter-feature correlations found in tabular data collected from real data sources” ( Jandial , p. 1, 0003). Jandial further teaches that the variational autoencoder is trained to learn inter-feature correlations from tabular data collected from real data sources and that synthetic tabular data produced by the generator model “maintains the inter-feature correlations present in the original tabular data collected from real data sources” ( Jandial , p. 1, 0004). Jandial also teaches that the synthetic tabular data generator generates records “that are realistic as compared to the collected tabular records 130, including preservation of inter-feature correlations within individual synthetic tabular data records and other characteristics exhibited by the collected tabular data records 130” ( Jandial , p. 3, 0026). Jandial further discloses that the collected tabular data has rows and columns/features, and that the learned feature relationships are used to generate synthetic records. In particular, Jandial states that “each row of the collected tabular data 205 define[s] a single tabular data record 260 and each of the columns 270 corresponds to an individual feature” ( Jandial , p. 4, 0031). Jandial also teaches that reconstruction-correlation loss is used to preserve inter-feature correlations, including correlations between columns of the original collected tabular records and the synthetic tabular records ( Jandial , p. 7, 0044–0045). Jandial further teaches that the method includes training the first machine-learning model using tabular records to learn correlations between features and training the second machine-learning model to generate synthetic tabular records based on those correlations ( Jandial , p. 8, 0049–0051). Thus, Jandial teaches that the synthetic structured table comprises synthetic structured data generated based on distributions/correlations determined from the real data table. Bruera and Jandial are analogous to the claimed invention because both are directed to generating synthetic data from real data for machine-learning applications while preserving useful statistical characteristics of the original data. Bruera teaches extracting structured attributes from real CV data, learning conditional dependencies and conditional probability distributions from those real samples using a Bayesian network, and generating synthetic candidate attributes from the learned distributions. Jandial teaches receiving real tabular data records, learning inter-feature correlations from those records, and generating synthetic tabular records based on the learned correlations. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to use Jandial ’s correlation-based synthetic tabular-data generation technique with Bruera ’s synthetic-data generation process to generate synthetic structured data based on distributions or correlations determined from the real data. The motivation to combine is to improve the realism and statistical fidelity of the synthetic structured data by preserving relationships, distributions, and correlations learned from the real data, as taught by Bruera and Jandial ( Bruera , p. 54, §1, left column, paragraph beginning “In our approach, we start from real samples…”; Jandial , p. 1, 0003–0004; Fig. 4, steps 410–414). Claims 8 and 15 are medium and system claims, respectively. Claims 8 and 15 contain similar limitations of claim 1. Therefore, claim 8 and 15 are rejected under the same rationale. As to dependent Claim 9 , the combination of Bruera , Jandial , and Hegselmann teaches the non-transitory computer-readable storage medium of Claim 8 as set forth above. Jandial teaches the computer-readable-storage-medium and processor-execution context inherited from Claim 8. Specifically, Jandial teaches that functions may be “carried out by hardware, firmware, software, and/or hardware executing firmware and/or software,” and that some functions may be carried out “by a processor executing instructions stored in memory as further described with reference to FIG. 9” ( Jandial , p. 3, 0022). Jandial further illustrates a computing device including “Memory 912” and “Processor(s) 914” ( Jandial , Fig. 9). Claim 9 further recites “wherein the prompt comprises rows of the sampled data table as few-shot examples for a LLM of the LLM system to generate the synthetic unstructured data.” Claim 9 recites the same substantive prompt limitation as Claim 2, except that Claim 9 is drafted in non-transitory computer-readable-storage-medium form rather than method form. Therefore, Claim 9 is rejected under the same rationale set forth above with respect to Claim 2. As to dependent Claims 10, 11, and 14 Claims 10, 11, and 14 depend from Claim 8 and are directed to a non-transitory computer-readable storage medium. Claims 10, 11, and 14 recite limitations corresponding to the limitations recited in Claims 3, 4, and 7, respectively, except that Claims 10, 11, and 14 are drafted in non-transitory computer-readable-storage-medium form rather than method form. Therefore, Claims 10, 11, and 14 are rejected under the same rationale set forth above with respect to Claims 3, 4, and 7. More specifically, Claim 10 is rejected under the same rationale set forth above with respect to Claim 3; Claim 11 is rejected under the same rationale set forth above with respect to Claim 4; Claim 14 is rejected under the same rationale set forth above with respect to Claim 7. As to dependent Claims 16, 17, and 18 Claims 16, 17, and 18 depend from Claim 15 and are directed to a system. Claims 16, 17, and 18 recite limitations corresponding to the limitations recited in Claims 2, 3, and 4, respectively, except that Claims 16-20 are drafted in system form rather than method form. Therefore, Claims 16, 17, and 18 are rejected under the same rationale set forth above with respect to Claims 2, 3, and 4. More specifically, Claim 16 is rejected under the same rationale set forth above with respect to Claim 2; Claim 17 is rejected under the same rationale set forth above with respect to Claim 3; Claim 18 is rejected under the same rationale set forth above with respect to Claim 4 . 07-21-aia AIA Claim 5, 12, and 19 are rejected under 35 U.S.C. §103 as being unpatentable over Bruera , Jandial , and Hegselmann , and further in view of Hazard et al. (Hazard) , U.S. Patent No. US 11,640,561 B2 . As to dependent Claim 5 , the combination of Bruera , Jandial , and Hegselmann teaches the method of Claim 1 as set forth above. Claim 5 further recites: “wherein providing an aggregate synthetic table comprises selectively filtering at least a portion of a semi-structured synthetic table that is provided from the LLM system.” Bruera teaches generating a semi-structured synthetic dataset from an LLM/generative language model. Specifically, Bruera teaches that its approach includes “the generation of a synthetic dataset of CVs,” wherein the final stage involves “sampling a synthetic set of attributes,” “inserting them in a series of ready-made prompts,” and “feeding these filled prompts, sequentially, to the generative language model so as to create a coherent CV” ( Bruera , p. 55, §3). Bruera further teaches that “[t]he candidate attributes sampled from the Bayesian network together with the artificial personal details are then used to generate the text for each section of the synthetic CV” ( Bruera , p. 56, §3.3). Bruera also teaches that the attributes are plugged into “linguistic structures called prompts” and presented “to the generative model as starting points for the generation of each section of the CV” ( Bruera , p. 56, §3.3). Thus, Bruera teaches a synthetic semi-structured dataset including structured attributes and generated unstructured text provided from a generative language model. Jandial teaches the synthetic tabular-data context for the aggregate synthetic table. Specifically, Jandial teaches generating synthetic tabular data records based on real tabular data and learned correlations, including “receiving a set of tabular data records” and training a machine-learning model “to generate a set of synthetic tabular data records” ( Jandial , front page, Abstract; Fig. 4, steps 410–414). Jandial further teaches executing a machine-learning model to generate “a set of synthetic tabular data records” and “stor[ing] the set of synthetic tabular data records” ( Jandial , Fig. 8, steps 810–814). Jandial also teaches that “synthesized tabular data records 132 can be used by application 103 to efficiently train machine learning model 105” ( Jandial , p. 3, 0026). However, the combination of Bruera , Jandial , and Hegselmann does not expressly teach selectively filtering at least a portion of the semi-structured synthetic table provided from the LLM system. Hazard et al. ( Hazard ) further teaches selectively filtering generated synthetic data. Specifically, Hazard teaches that generated synthetic data may be checked for similarity against training data and then “resampled, removed, and/or replaced” if similarity conditions are met ( Hazard , front page, Abstract). Hazard also teaches a process including “Test synthetic data case for fitness and/or similarity” followed by “Filter data” ( Hazard , Fig. 1D, steps 160, 165). Hazard further explains that “FIG. 1D is a flow diagram depicting example processes for synthetic data generation in computer-based reasoning systems” and that “the process 100 depicted in FIG. 1D includes filtering 165 data from the generated synthetic data” ( Hazard , p. 28). Hazard states that the process may include generating more synthetic data than needed and “then filtering 165 the synthetic data back down to a desired size (or size range) based at least in part on the one or more hypergraph feature filters” ( Hazard , p. 28). Hazard further teaches that, in some embodiments, “it may be useful to generate not only the types of hypergraph features that will later be filtered, but also types of hypergraph features that will not be later filtered,” and provides examples of generating more data than needed and filtering based on desired features ( Hazard , p. 30). Hazard also teaches threshold-based quality control for synthetic data. Specifically, Hazard teaches determining synthetic data based on training data cases, determining a dataset quality metric, determining whether the dataset quality metric is beyond a threshold, taking corrective action, and providing synthetic data for use in a computer-based reasoning model ( Hazard , Fig. 6, steps 620–650, 170). Hazard explains that determining a dataset quality metric may include statistical quality metrics, model comparison metrics, and privacy metrics ( Hazard , pp. 31–32). Hazard further teaches that statistical quality metrics compare statistical properties of training data cases and synthetic data cases, model comparison metrics compare machine-learning model performance, and privacy metrics measure privacy or similarity of data between two datasets ( Hazard , p. 32). Thus, Hazard teaches selectively filtering generated synthetic data based on fitness, similarity, feature filters, and quality metrics. Bruera , Jandial , and Hazard are analogous to the claimed invention because each relates to synthetic-data generation or evaluation for machine-learning or computer-based reasoning systems. Bruera teaches generating semi-structured synthetic CV data using structured attributes and a generative language model. Jandial teaches generating synthetic tabular data records from real tabular data for machine-learning applications. Hazard teaches filtering and quality control for generated synthetic data before providing the synthetic data for use in a computer-based reasoning model. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the synthetic-data generation process of Bruera and Jandial to include Hazard ’s selective filtering of generated synthetic data. The motivation to combine is provided by Hazard , which teaches testing generated synthetic data for fitness and/or similarity, filtering generated synthetic data, and applying quality metrics and thresholds before providing the synthetic data for model use ( Hazard , Fig. 1D, steps 160–170; Fig. 6, steps 620–650; pp. 28, 31–32). A person of ordinary skill would have been motivated to apply such filtering to improve the quality, suitability, and privacy-related characteristics of the aggregate synthetic table before using it for machine-learning training. Claim 12 depends from Claim 8 and is directed to a non-transitory computer-readable storage medium. Claim 12 recites a limitation is the same as the limitation recited in Claim 5, except that Claim 12 is drafted in non-transitory computer-readable-storage-medium form rather than method form. Therefore, Claim 12 is rejected under the same rationale set forth above with respect to Claim 5. Claim 19 depends from Claim 15 and is directed to a system. Claim 19 recites the limitation is the same as the limitation recited in Claim 5, except that Claim 19 is drafted in system form rather than method form. Therefore, Claim 19 is rejected under the same rationale set forth above with respect to Claim 5 . 07-21-aia AIA Claim 6, 13, and 20 are rejected under 35 U.S.C. §103 as being unpatentable over Bruera , Jandial , and Hegselmann , and further in view of Sankaranarayanan et al. (Sankaranarayanan) , U.S. Patent No. US 11,687,568 B2 . As to dependent Claim 6 , the combination of Bruera , Jandial , and Hegselmann teaches the method of Claim 1 as set forth above. Claim 6 further recites: “wherein providing an aggregate synthetic table comprises aggregating at least portions of multiple semi-structured synthetic table.” Bruera teaches generating semi-structured synthetic data from structured attributes and generated natural-language text. Specifically, Bruera teaches that its approach includes “the generation of a synthetic dataset of CVs,” wherein the final stage includes “sampling a synthetic set of attributes,” “inserting them in a series of ready-made prompts,” and “feeding these filled prompts, sequentially, to the generative language model so as to create a coherent CV” ( Bruera , p. 55, §3). Bruera further teaches that “[t]he candidate attributes sampled from the Bayesian network together with the artificial personal details are then used to generate the text for each section of the synthetic CV” ( Bruera , p. 56, §3.3). Bruera also teaches that the authors “generate in this way a set of 4000 synthetic CVs” ( Bruera , p. 58). Thus, Bruera teaches generating multiple semi-structured synthetic records including structured attributes and generated unstructured text. Jandial teaches generating and storing sets of synthetic tabular data records. Specifically, Jandial teaches generating synthetic tabular data by “receiving a set of tabular data records” and training a machine-learning model “to generate a set of synthetic tabular data records” ( Jandial , front page, Abstract; Fig. 4, steps 410–414). Jandial further teaches executing a machine-learning model to “generate a set of synthetic tabular data records” and “store the set of synthetic tabular data records” ( Jandial , Fig. 8, steps 810–814). Jandial also teaches that “synthesized tabular data records 132 can be used by application 103 to efficiently train machine learning model 105” ( Jandial , p. 3, 0026). Thus, Jandial supports the synthetic tabular dataset context for the claimed aggregate synthetic table. However, the combination of Bruera , Jandial , and Hegselmann does not expressly teach that providing the aggregate synthetic table comprises aggregating at least portions of multiple semi-structured synthetic tables. Sankaranarayanan further teaches aggregating or augmenting dataset portions with generated synthetic data to create a new synthetic dataset. Specifically, Sankaranarayanan teaches that a data catalog system may “automatically generate synthetic datasets based upon original datasets” and that the synthetic dataset may comprise “synthetic data that is generated using one or more data generation techniques” ( Sankaranarayanan , front page, Abstract). Sankaranarayanan further teaches generating and storing “a new synthetic dataset based upon the original dataset” using one or more machine-learning techniques ( Sankaranarayanan , Fig. 2, step 210; front page, Abstract). Sankaranarayanan teaches one aggregation technique in Figure 4. Figure 4 recites “copy the original dataset to generate a new dataset,” “generate new synthetic data for the missing data,” “augment the new dataset generated in 404 with the new synthetic data generated in 406 to create a new synthetic dataset,” and “store the new synthetic dataset generated in 408” ( Sankaranarayanan , Fig. 4, steps 404–410). The specification corresponding to Figure 4 likewise states that “At 406, new synthetic data is generated,” ( Sankaranarayanan , p. 29, column 23, para. 1) that “At 408, the new dataset generated in 404 is augmented with the new synthetic data generated in 406 to create a new synthetic dataset,” ( Sankaranarayanan , p. 29, column 23, para. 2) and that “At 410, the new synthetic dataset generated in 408 is stored” ( Sankaranarayanan , p. 29, column 23, para. 3). This teaches aggregating generated synthetic data into another dataset to form a new synthetic dataset. Sankaranarayanan also teaches replacing or substituting portions of a dataset with synthetic data. Figure 5 recites identifying portions of an original dataset containing restricted data, copying the original dataset to generate a new dataset, generating new synthetic data for the restricted data, replacing or substituting the restricted data portions with the new synthetic data to create a new synthetic dataset, and storing the new synthetic dataset ( Sankaranarayanan , Fig. 5, steps 502–510). The specification similarly states that “new synthetic data is generated based upon the original dataset” and that “At 508, the new synthetic data generated in 506 replaces/substitutes the restricted data in the new dataset copied in 504 to create a new synthetic dataset” ( Sankaranarayanan , p. 29, column 24, para. 3). This supports aggregating retained portions of a dataset with synthetic replacement portions. Sankaranarayanan further teaches augmenting a copied dataset with newly generated synthetic data. Figure 6B recites “copy the original dataset to generate a new dataset,” “generate new synthetic data for augmenting the new dataset with synthetically generated data,” “add the new synthetic data generated in step 614 to the new dataset generated in 612 to create a new synthetic dataset,” and “store the new synthetic dataset generated in 616” ( Sankaranarayanan , Fig. 6B, steps 612–618). The specification corresponding to Figure 6B states that “At 614, new synthetic data is generated based upon the original dataset identified in 402,” that the generated synthetic data may “augment portions of the generated new dataset,” ( Sankaranarayanan , p. 30, column 26, para. 2) and that “At 616, the new dataset generated in 612 is augmented by adding the new synthetic data generated in 614 to the new dataset to create a new synthetic dataset” ( Sankaranarayanan , p. 30, column 26, para. 3). Thus, Sankaranarayanan teaches aggregating at least portions of multiple generated or copied dataset components to create a new synthetic dataset. Bruera , Jandial , and Sankaranarayanan are analogous to the claimed invention because each relates to generating, organizing, or using synthetic data for machine-learning or data-processing applications. Bruera teaches generating semi-structured synthetic CV records using structured attributes and generated text. Jandial teaches generating synthetic tabular data records from real tabular data for machine-learning applications. Sankaranarayanan teaches augmenting, replacing, and adding generated synthetic data into datasets to create and store new synthetic datasets. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the synthetic-data generation process of Bruera and Jandial to include Sankaranarayanan ’s dataset augmentation and aggregation techniques. The motivation to combine is to create a usable aggregate synthetic dataset by adding, replacing, or augmenting dataset portions with generated synthetic data, as taught by Sankaranarayanan , which discloses augmenting a dataset with new synthetic data to create a new synthetic dataset and storing the resulting synthetic dataset ( Sankaranarayanan , Fig. 4, steps 406–410; Fig. 6B, steps 614–618; column. 23–26). A person of ordinary skill would have been motivated to apply such aggregation to the semi-structured synthetic data generated by Bruera and the synthetic tabular records generated by Jandial to form an aggregate synthetic table suitable for downstream machine-learning use. Claim 13 depends from Claim 8 and is directed to a non-transitory computer-readable storage medium. Claim 13 recites a limitation is the same as the limitation recited in Claim 6, except that Claim 13 is drafted in non-transitory computer-readable-storage-medium form rather than method form. Therefore, Claim 13 is rejected under the same rationale set forth above with respect to Claim 6. Claim 20 depends from Claim 15 and is directed to a system. Claim 20 recites a limitation is the same as the limitation recited in Claim 6, except that Claim 20 is drafted in system form rather than method form. Therefore, Claim 20 is rejected under the same rationale set forth above with respect to Claim 6. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to HUNG VAN LE whose telephone number is (571)270-0164. The examiner can normally be reached 8 a.m. - 5 p.m.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HUNG VAN LE/Examiner, Art Unit 2145 /CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145 Application/Control Number: 18/455,775 Page 2 Art Unit: 2145 Application/Control Number: 18/455,775 Page 3 Art Unit: 2145 Application/Control Number: 18/455,775 Page 4 Art Unit: 2145 Application/Control Number: 18/455,775 Page 5 Art Unit: 2145 Application/Control Number: 18/455,775 Page 6 Art Unit: 2145 Application/Control Number: 18/455,775 Page 7 Art Unit: 2145 Application/Control Number: 18/455,775 Page 8 Art Unit: 2145 Application/Control Number: 18/455,775 Page 9 Art Unit: 2145 Application/Control Number: 18/455,775 Page 10 Art Unit: 2145 Application/Control Number: 18/455,775 Page 11 Art Unit: 2145 Application/Control Number: 18/455,775 Page 12 Art Unit: 2145 Application/Control Number: 18/455,775 Page 13 Art Unit: 2145 Application/Control Number: 18/455,775 Page 14 Art Unit: 2145 Application/Control Number: 18/455,775 Page 15 Art Unit: 2145 Application/Control Number: 18/455,775 Page 16 Art Unit: 2145 Application/Control Number: 18/455,775 Page 17 Art Unit: 2145 Application/Control Number: 18/455,775 Page 18 Art Unit: 2145 Application/Control Number: 18/455,775 Page 19 Art Unit: 2145 Application/Control Number: 18/455,775 Page 20 Art Unit: 2145 Application/Control Number: 18/455,775 Page 21 Art Unit: 2145 Application/Control Number: 18/455,775 Page 22 Art Unit: 2145 Application/Control Number: 18/455,775 Page 23 Art Unit: 2145 Application/Control Number: 18/455,775 Page 24 Art Unit: 2145 Application/Control Number: 18/455,775 Page 25 Art Unit: 2145
Read full office action

Prosecution Timeline

Aug 25, 2023
Application Filed
May 29, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month