Prosecution Insights
Last updated: October 01, 2026
Application No. 19/345,953

DATASET PREPARATION

Non-Final OA §101§103§112
Filed
Sep 30, 2025
Priority
Aug 31, 2023 — continuation of 12/450,289
Examiner
CAO, PHUONG THAO
Art Unit
2164
Tech Center
2100 — Computer Architecture & Software
Assignee
International Business Machines Corporation
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
1y 11m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
608 granted / 778 resolved
+23.1% vs TC avg
Moderate +15% lift
Without
With
+14.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
9 currently pending
Career history
799
Total Applications
across all art units

Statute-Specific Performance

§101
17.6%
-22.4% vs TC avg
§103
41.4%
+1.4% vs TC avg
§102
6.6%
-33.4% vs TC avg
§112
24.9%
-15.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 778 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to Application filed on 09/30/2025. Claims 1-20 are pending. Priority This application is claimed as a continuation of U.S. Patent Application No. 18/240,866 filed on 08/31/2023. However, the parent application does not provide support for at least limitations “column-to-row association” and “row-to-row association” as recited in independent claim 1, at least limitation “a user objective specification defining a target analytical task to be supported by a production dataset” as recited in independent claim 8, and at least limitations “versioned production datasets” and “predefined selection rule” as recited in independent claim 15. Therefore, all claims 1-20 of the claimed invention of this application are not entitled to the effective filing date of the parent application. Information Disclosure Statement The Information Disclosure Statement filed by Applicant on 11/20/2025 has been considered. A copy of the considered IDS is enclosed with this Office action. Specification The disclosure is objected to because of the following informalities: Regarding paragraph [0001] of the Specification, the U.S. Patent Application No. 18/240,866 has been patented, its information should be supplemented with its patent information (e.g., Patent No.). Appropriate correction is required. The specification is objected to as failing to provide proper antecedent basis for the claimed subject matter. See 37 CFR 1.75(d)(1) and MPEP § 608.01(o). Correction of the following is required: the following recited terms/limitations, which are not even mentioned by the Specification, include “column-to-row associations” and “row-to-row associations” (see line 3 of claim 1), “row labels” (see line 2 of claim 3), “a user objective specification defining a target analytical task” (see line 2 of claim 8), “versioned production datasets”, “a respective composite dataset score” and “predefined selection rule” (see claim 15). Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 1, it is unclear regarding the meaning of “column-to-row associations” or “row-to-row associations” as recited in the claim without any definition and/or disclosure in the Specification. Therefore, the metes and bounds of the claimed invention are indefinite. Claim 2 recites the limitation "the associated elements" in line 2. There is insufficient antecedent basis for this limitation in the claim. In addition, it is unclear what is considered as “an element” as recited. It should be noted that the Specification only discloses a column association strength value between any first and second columns (see [0053]). Regarding claim 3, it is unclear what are considered as “row labels” as claimed without defining/disclosing in the Specification. In addition, the Specification only discloses performing semantic similarity based on column matching or calculating semantic similarity score for identifying column matching and/or identifying dataset for merging in response to user search input, but does not disclose storing/including semantic similarity scores in the index metadata as recited (see Specification, [0071]-[0074]). Regarding claim 4, it is unclear how ranking candidate dataset groups for merging is in dependence on aggregated strength values across the column-to-column, column-to-row, and row-to-row associations as recited since the Specification does not mention “column-to-row association” or “row-to-row association”, let alone any strength value associated with a column-to-row association or a row-to-row association. The Specification only discloses ranking candidate dataset groups based a set of factors and their associated weight (see Specification, [0080]-[0085]). Claim 5 recites the limitation "the datasets" in line 2 and the limitation “the metadata” in line 2. There is insufficient antecedent basis for these limitations in the claim. Regarding claim 6, it is unclear what is considered as “a minimum threshold association strength across at least two types of associations” as recited. The Specification only mentions a strength value for a column association and various types of column associations, e.g., string-string, int-int, string-int, etc. (see Specification, [0052]). Dependent claim 7 is rejected for incorporating and failing to resolve the deficiency of the rejected independent claim 1 upon which it depends. Claim 8 recites the limitation "the candidate datasets" in line 5 and line 6. There is insufficient antecedent basis for this limitation in the claim. In addition, it is unclear regarding limitation “a user objective specification defining a target analytical task” or “a target analytical task to be supported by a production dataset”, which is recited but is not defined or disclosed in the Specification. Therefore, the metes and bounds of the claimed invention are indefinite. Regarding claims 12-14, instances of limitation “the candidate datasets” recited in these claims should be amended in accordance to claim 8. Other dependent claims 9-14 are rejected as incorporating and failing to resolve the deficiency of the rejected independent claim 8 upon which they depend. Regarding claim 15, it is unclear regarding “versioned production datasets” and “a respective composite dataset score” as recited without definition or disclosure in the Specification. The specification only discloses versions of candidate group(s) of datasets for merging and presenting the versions of candidate groups along with a rank/score and prompting data to a user for selection groups of datasets for merging (see Specification, [0086]-[0088]). The specification does not define or disclose version production datasets or a respective composite dataset score for each version production dataset as recited. In addition, it is unclear regarding “an automated selection…based on a user choice”, which raises questions (e.g., why an automated selection is based on a user choice (?)). Therefore, the metes and bounds of the claimed invention are indefinite. Other dependent claims 16-20 are rejected as incorporating and failing to resolve the deficiency of the rejected independent claim 15 upon which they depend. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of analyzing data without significantly more. The claims recite an abstract idea of generating metadata for datasets, merging datasets and generating a production datasets based on metadata and/or associations, which are broadly recited steps/concepts that can be performed in the human mind or with the aid of pencil and paper and directed to mental processes grouping of abstract ideas . This judicial exception is not integrated into a practical application because other additional elements including common/routine computer functionality (e.g., accessing, storing, displaying, etc.) and/or insignificant extra-solution activity (e.g., mere data gathering and displaying) for implementing the abstract idea are not sufficient to integrate the abstract idea into a practical application. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because additional elements include common/routine computer functions (e.g., accessing, storing, displaying, etc.) and/or insignificant extra-solution activity (e.g., mere data gathering and displaying), which are not sufficient to amount to significantly more than the recited abstract idea. Abstract idea analysis as follows: Step 1: According to the first part of the analysis, in the instant claims, claims 1-20 are directed to a computer implemented method comprising a series of steps (i.e. a process). Thus, each of the claims falls within one of the four statutory categories (i.e. process, machine, manufacture or composition of matter). Step 2a Prong 1 (claims 1, 8 and 15): Regarding independent claim 1, the following limitations recited in claim 1 are abstract ideas that fall under mental processes: generating, for a plurality of datasets, multi-dimensional index metadata that specifies column-to-column associations, column-to-row associations, and row-to-row associations, each association being quantified by a respective strength value (the step of generating multi-dimensional index metadata as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper through mental processes (e.g., observation, evaluation, judgment and opinion), for instance, a human can observe a plurality of datasets to recognize associations between data/structure elements and assign any appropriate value for each association), and producing a production dataset by selecting and merging portions of the plurality of datasets in dependence on the strength values across at least two different types of associations specified in the multi-dimensional index metadata (the step of producing a production dataset as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper, e.g., a human mind can think/plan on creating a database based on merging/combining data from a plurality of datasets based on associations regarding data or data structure). All the limitations above are mental steps that can be performed in the human mind or with the aid of pencil and paper. Regarding independent claim 8, the following limitations recited in claim 8 are abstract ideas that fall under mental processes: examining a plurality of candidate datasets in dependence on metadata indices describing associations among columns within the candidate datasets (the step of examining a plurality of candidate datasets as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper through mental processes (e.g., observation, evaluation, judgment and opinion), for instance, a human can observe a plurality of datasets to examine and/or recognize associations among columns within the dataset); and merging the candidate datasets into the production dataset in a manner that varies according to the received user objective specification (this step of merging datasets to generate a production dataset as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper, for instance, a human mind can think or plan on how to merge dataset to generate a new dataset or simply illustrate on a paper the merging/combining the datasets). All the limitations above are mental steps that can be performed in the human mind or with the aid of pencil and paper. Regarding independent claim 15, the following limitations recited in claim 15 are abstract ideas that fall under mental processes: generating a plurality of versioned production datasets, each version produced by merging different subsets of columns from a plurality of source datasets (this step of generating as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper, for instance, the human mind can think or plan on different way to combine/merge datasets); computing a respective composite dataset score for each versioned production dataset in dependence on metadata indices that quantify associations between the merged columns (this step of computing as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper, for instance, the human mind can do a simple calculation and/or assignment of values/scores); and initiating an automated selection, deployment, or downstream processing of a versioned production dataset based on a user choice or predefined selection rule (this step of initiating as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper, for instance, the human mind can think about how to select, process or use a dataset). All the limitations above are mental steps that can be performed in the human mind or with the aid of pencil and paper. Step 2a Prong 2 (Claims 1, 8 and 15): Regarding claim 1, the following limitations in claim 1 are additional elements: storing the multi-dimensional index metadata in memory (the step of storing as broadly recited is directed to generic computer component (e.g., memory) for performing common and routine computer function (e.g., storing data in memory), which is directed to insignificant extra-solution activity). These are generic computer components, common and routing computer functions or insignificant extra-solution activity for implementing or applying the abstract. Accordingly, these additional elements do not integrate the abstract idea(s) into a practical application because they do not impose any meaningful limits on practicing the abstract idea(s). Regarding claim 8, the following limitations in claim 8 are additional elements: receiving a user objective specification defining a target analytical task to be supported by a production dataset (this step of receiving as broadly recited is directed to data gathering recited at high level of generality and being insignificant extra-solution activity). These are insignificant extra-solution activity for implementing or applying the abstract. Accordingly, these additional elements do not integrate the abstract idea(s) into a practical application because they do not impose any meaningful limits on practicing the abstract idea(s). Regarding claim 15, the following limitations in claim 15 are additional elements: presenting the plurality of versioned production datasets and their respective composite dataset scores for display in a user interface (this step of presenting as broadly recited is directed to data outputting/displaying recited at high level of generality and being insignificant extra-solution activity). These are insignificant extra-solution activity for implementing or applying the abstract. Accordingly, these additional elements do not integrate the abstract idea(s) into a practical application because they do not impose any meaningful limits on practicing the abstract idea(s). Step 2b (Claims 1, 8 and 15): Regarding claim 1, the following limitations in claim 1 are additional elements: storing the multi-dimensional index metadata in memory (the step of storing as broadly recited is directed to generic computer component (e.g., memory) for performing common and routine computer function (e.g., storing data in memory), which is directed to insignificant extra-solution activity). These are generic computer components, common and routing computer functions or insignificant extra-solution activity or well-understood, routine, conventional activity, and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 8, the following limitations in claim 8 are additional elements: receiving a user objective specification defining a target analytical task to be supported by a production dataset (this step of receiving as broadly recited is directed to data gathering recited at high level of generality and being insignificant extra-solution activity). These are insignificant extra-solution activity or well-understood, routine, conventional activity, and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 15, the following limitations in claim 15 are additional elements: presenting the plurality of versioned production datasets and their respective composite dataset scores for display in a user interface (this step of presenting as broadly recited is directed to data outputting/displaying recited at high level of generality and being insignificant extra-solution activity). These are insignificant extra-solution activity or well-understood, routine, conventional activity, and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 2, claim 2 depends on claim 1. As such, claim 2 recites the abstract idea as presented in claim 1. In addition, claim 2 includes additional elements: wherein the strength value of an association is determined in dependence on at least one factor selected from: a co-occurrence frequency of the associated elements across datasets, a predictive accuracy of inferring values of one element from another, and a derivability measure indicating whether a value of one element can be derived from another (this element recited determining the strength value of an association, which is broadly recited without specifying “how” and can be mentally performed in a human mind, e.g., just thinking about it without actually performing the calculation). These are additional elements directed to mental process or abstract idea, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 3, claim 3 depends on claim 1. As such, claim 3 recites the abstract idea as presented in claim 1. In addition, claim 3 includes additional elements: wherein the multi-dimensional index metadata further includes semantic similarity scores among column names, row labels, or data values determined by natural language processing (this element specifying data regarding the multi-dimensional index metadata without specifying any functionality, which is directed to additional data). These are additional elements directed to additional data/information, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 4, claim 4 depends on claim 1. As such, claim 4 recites the abstract idea as presented in claim 1. In addition, claim 4 includes additional elements: wherein producing the production dataset comprises ranking candidate dataset groups for merging in dependence on aggregated strength values across the column-to-column, column-to-row, and row-to-row associations (this element including a ranking step that is recited broadly and can be mentally performed in the human mind or with the aid of pencil and paper, for instance, the human mind can ranking data by observing and assigning ranking). These are additional elements directed to mental process or abstract idea, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 5, claim 5 depends on claim 1. As such, claim 5 recites the abstract idea as presented in claim 1. In addition, claim 5 includes additional elements: wherein the storing comprises persisting the multi-dimensional index metadata in a memory location separate from the datasets and updating the metadata in response to ingestion of new datasets (the steps of storing as recited broadly is directed to common/routine computer function (e.g., storing data in memory) and the step of updating as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper, for instance, the human mind can update/modify data/dataset or knowledge upon receiving new data/dataset). These are additional elements directed to generic computer component, common/routine computer function, and mental process or abstract idea, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 6, claim 6 depends on claim 1. As such, claim 6 recites the abstract idea as presented in claim 1. In addition, claim 6 includes additional elements: wherein producing the production dataset comprises filtering out candidate dataset groups that fail to satisfy a minimum threshold association strength across at least two types of associations (the step of filtering out as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper). These are additional elements directed to mental process or abstract idea, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 7, claim 7 depends on claim 1. As such, claim 7 recites the abstract idea as presented in claim 1. In addition, claim 7 includes additional elements: wherein producing the production dataset comprises performing dynamic semantic merging that resolves synonymous or abbreviated column names prior to merging portions of the plurality of datasets (the step of merging as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper). These are additional elements directed to mental process or abstract idea, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 9, claim 9 depends on claim 8. As such, claim 9 recites the abstract idea as presented in claim 8. In addition, claim 9 includes additional elements: wherein the user objective specification includes at least one constraint selected from: a required set of columns, a semantic tag, a column value range, and a sample count (this element specifying the user objective specification without reciting any functionality, which is directed to mere additional data). These are mere additional data or information, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 10, claim 10 depends on claim 8. As such, claim 10 recites the abstract idea as presented in claim 8. In addition, claim 10 includes additional elements: further comprising presenting a ranked list of candidate dataset groups satisfying the user objective specification on a user interface for user selection (this step of presenting as broadly recited is directed to mere data outputting/displaying recited at high level of generality and being insignificant extra-solution activity). These are insignificant extra-solution activity or well-understood, routine and conventional activity, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 11, claim 11 depends on claim 8. As such, claim 11 recites the abstract idea as presented in claim 8. In addition, claim 11 includes additional elements: wherein the merging comprises selecting a merge strategy from among a plurality of available merge strategies, the merge strategy being selected in accordance with the user objective specification (this element specifying the merging to comprise selecting step, which as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper, for instance, providing a set of merge strategies, the human mind can decide on which merge strategy to select/use). These are directed to mental steps/process or abstract idea, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 12, claim 12 depends on claim 8. As such, claim 12 recites the abstract idea as presented in claim 8. In addition, claim 12 includes additional elements: wherein examining the candidate datasets comprises performing semantic similarity matching between terms in the user objective specification and column names of the candidate datasets (this element specifying the examining to comprise performing semantic similarity matching step, which as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper, for instance, the human mind can observe and recognize the semantic similarity matching). These are directed to mental steps/process or abstract idea, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 13, claim 13 depends on claim 8. As such, claim 13 recites the abstract idea as presented in claim 8. In addition, claim 13 includes additional elements: wherein examining the candidate datasets comprises performing semantic similarity matching between terms in the user objective specification and column names of the candidate datasets (this element specifying the examining to comprise performing semantic similarity matching step, which as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper, for instance, the human mind can observe and recognize the semantic similarity matching), and wherein the semantic similarity matching is performed using at least one of: word embedding vector distances, N-gram comparisons, and abbreviation expansion (this element specifying the techniques/methods used without specifying “how” (i.e., the process of acts/functions) is directed to mere additional data). These are directed to mental steps/process or abstract idea and mere additional data, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 14, claim 14 depends on claim 8. As such, claim 14 recites the abstract idea as presented in claim 8. In addition, claim 14 includes additional elements: automatically merging the candidate datasets without user confirmation in response to receiving a predefined user objective specification (the merging step as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper, wherein “automatically” is directed to generic computer and/or computer components for implementing the mental step or abstract idea). These are directed to generic computer and/or computer components for implementing or applying mental steps/process or abstract idea, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 16, claim 16 depends on claim 15. As such, claim 16 recites the abstract idea as presented in claim 15. In addition, claim 16 includes additional elements: wherein computing the composite dataset score comprises applying weighted factors including at least one of: a common column factor, an association strength factor, and a semantic similarity factor (the computing step comprising applying as broadly recited without specifying “how” can be mentally performed in the human mind or with the aid of pencil and paper, for instance, the human mind can perform a simple calculating based on many factors/parameters). These are directed to mental steps/process or abstract idea, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 17, claim 17 depends on claim 15. As such, claim 17 recites the abstract idea as presented in claim 15. In addition, claim 17 includes additional elements: wherein presenting the versioned production datasets comprises ranking the versioned production datasets in accordance with the composite dataset scores (the step of ranking as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper, for instance, the human mind can perform a ranking based on observation, evaluation, judgment and opinion). These are directed to mental steps/process or abstract idea, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 18, claim 18 depends on claim 15. As such, claim 18 recites the abstract idea as presented in claim 15. In addition, claim 18 includes additional elements: wherein initiating the automated selection comprises deploying the selected versioned production dataset for training a machine learning model (the step of initiating or deploying as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper, for instance, the human mind can determine on how to use/provide data/dataset, in addition, the step of deploying as broadly recited is directed to mere data gathering or outputting recited at high level of generality or insignificant extra-solution activity). These are directed to mental steps/process or abstract idea and insignificant extra-solution activity for implementing or applying the abstract idea, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 19, claim 19 depends on claim 15. As such, claim 19 recites the abstract idea as presented in claim 15. In addition, claim 19 includes additional elements: wherein initiating the automated selection comprises deploying the selected versioned production dataset for testing an application programming interface (API) (the step of initiating or deploying as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper, for instance, the human mind can determine on how to use/provide data/dataset, in addition, the step of deploying as broadly recited is directed to mere data gathering or outputting recited at high level of generality or insignificant extra-solution activity). These are directed to mental steps/process or abstract idea and insignificant extra-solution activity for implementing or applying the abstract idea, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Regarding claim 20, claim 20 depends on claim 15. As such, claim 20 recites the abstract idea as presented in claim 15. In addition, claim 20 includes additional elements: wherein initiating the automated selection comprises controlling an enterprise process in dependence on predictions generated by a machine learning model trained using the selected versioned production dataset (the step of initiating or controlling as broadly recited can be mentally performed in the human mind or with the aid of pencil and paper). These are directed to mental steps/process or abstract idea, which do not integrate the judicial exception into a practical application and do not amount to significantly more, see MPEP 2106.05(d)(II). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-7 (effective filing date 09/30/2025), as best understood, are rejected under 35 U.S.C. 103 as being unpatentable over Oberbreckling et al. (U.S. Patent No. 10,650,000, Patent date 05/12/2020) and further in view of Cronin et al. (U.S. Patent No. 9,367,853, Patent date 06/14/2016). As to claim 1, Oberbreckling et al. teaches: “A computer implemented method” (see Oberbreckling et al., Abstract) comprising: “generating, for a plurality of datasets, multi-dimensional index metadata that specifies column-to-column associations, column-to-row associations, and row-to-row associations, each association being quantified by a respective strength value” (see Oberbreckling et al., [column 7, lines 12-20], [column 23, line 52 to column 24, line 10] and [column 24, lines 55-67] for processing data/datasets from different data sources to identify/generate metadata or profile information (i.e., information about datasets), wherein metadata including relationships between datasets, wherein the relationships include column pairs (i.e., column-to-column associations), wherein each column pair is associated with a pair score; also see [column 30, line 66 to column 31, line 10); “storing the multi-dimensional index metadata in memory” (see Oberbreckling et al., [column 6, lines 43-48] for storing metadata in the distributed storage system or in a separate repository accessible to the data enrichment service); and “producing a production dataset by selecting and merging portions of the plurality of datasets in dependence on the strength values across at least two different types of associations specified in the multi-dimensional index metadata” (see Oberbreckling et al., [column 31, lines 20-44] for generating the new dataset as a merger of the datasets based on selected column pair (i.e., column-to-column association); also see [column 31, lines 2-9] wherein each column pair or relationship is associated with a score to reflect the strength of the relationship; also see [column 30, lines 53-65] wherein the pair score associated with each column pair may be determined based on a summary of the plurality of weighted scores; also see [column 32, lines 43-56] wherein join/merge functions are also based on row matching (i.e., row-to-row relationship/association)). However, Oberbreckling et al. does not explicitly teach metadata of datasets including column-to-row or row-to-column associations/relationships as recited. On the other hand, Cronin et al. explicitly teaches metadata of datasets including column-to-row or row-to-column associations/relationships (see Cronin et al., Abstract and [column 28, lines 23-40 for generating metadata for dataset(s) including indices representing probabilistic relationships between the rows and the columns of the dataset). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Cronin et al.'s teaching to Oberbreckling et al.’s system by implementing a feature for generating indices representing the relationships between rows and columns of a dataset. Skilled artisan would have been motivated to do so to provide Oberbreckling et al.’s system with an effective way to access datasets using metadata/indices. In addition, both of the references (Oberbreckling et al. and Cronin et al.) teach features that are directed to analogous art and they are directed to the same field of endeavor, such as, processing datasets to generate data/metadata associated with datasets including associations/relationship and forming/generating a dataset by merging/joining datasets. This close relation between both of the references highly suggests an expectation of success when combined. As to claim 2, this claim is rejected based on the same arguments as above to reject claim 1 and is similarly rejected including the following: Oberbreckling et al. as modified by Cronin et al. teaches: “wherein the strength value of an association is determined in dependence on at least one factor selected from: a co-occurrence frequency of the associated elements across datasets, a predictive accuracy of inferring values of one element from another, and a derivability measure indicating whether a value of one element can be derived from another” (see Oberbreckling et al., [column 16, line 56 to column 17, line 1] for a pattern metric (e.g., a statistical frequency of different patterns in the data), wherein each pattern in the data/dataset represents an element in the data/dataset). As to claim 3, this claim is rejected based on the same arguments as above to reject claim 1 and is similarly rejected including the following: Oberbreckling et al. as modified by Cronin et al. teaches: “wherein the multi-dimensional index metadata further includes semantic similarity scores among column names, row labels, or data values determined by natural language processing” (see Oberbreckling et al., [column 20, lines 25-50] for determining the semantic similarity between two or more datasets; also see [column 21, lines 14-28] for similarity metric/score). As to claim 4, this claim is rejected based on the same arguments as above to reject claim 1 and is similarly rejected including the following: Oberbreckling et al. as modified by Cronin et al. teaches: “wherein producing the production dataset comprises ranking candidate dataset groups for merging in dependence on aggregated strength values across the column-to-column, column-to-row, and row-to-row associations” (see Oberbreckling et al., [column 30, line 66 to column 31, line 9] for ranking column pairs, wherein each column pair representing a suggestion for a possible join of the datasets can be interpreted as equivalent to a candidate dataset group as recited). As to claim 5, this claim is rejected based on the same arguments as above to reject claim 1 and is similarly rejected including the following: Oberbreckling et al. as modified by Cronin et al. teaches: “wherein the storing comprises persisting the multi-dimensional index metadata in a memory location separate from the datasets and updating the metadata in response to ingestion of new datasets” (see Oberbreckling et al., [column 6, lines 43-48] for storing metadata in a distributed storage system or in a separate repository accessible to the data enrichment service). As to claim 6, this claim is rejected based on the same arguments as above to reject claim 1 and is similarly rejected including the following: Oberbreckling et al. as modified by Cronin et al. teaches: “wherein producing the production dataset comprises filtering out candidate dataset groups that fail to satisfy a minimum threshold association strength across at least two types of associations” (see Oberbreckling et al., [column 34, lines 34-47] for filtering one or more column pairs based on threshold). As to claim 7, this claim is rejected based on the same arguments as above to reject claim 1 and is similarly rejected including the following: Oberbreckling et al. as modified by Cronin et al. teaches: “wherein producing the production dataset comprises performing dynamic semantic merging that resolves synonymous or abbreviated column names prior to merging portions of the plurality of datasets” (see Oberbreckling et al., [column 35, lines 5-12] for determining semantic categories or attributes; also see [column 23, lines 5-17] for performing semantic similarity between two datasets) Claims 8-20 (effective filing date 09/30/2025), as best understood, are rejected under 35 U.S.C. 103 as being unpatentable over Oberbreckling et al. (U.S. Patent No. 10,650,000, Patent date 05/12/2020) and further in view of Dugan et al. (U.S. Patent No. 11,599,539, Patent date 03/07/2023). As to claim 8, Oberbreckling et al. teaches: “A computer implemented method” (see Oberbreckling et al., Abstract) comprising: “receiving a user objective specification defining a target analytical task to be supported by a production dataset” (see Oberbreckling et al., [column 11, lines 40-52] for receiving a data enrichment request from the client (i.e., a user), the data enrichment request identifies a data source and/or particular data (e.g., tables, columns or other data available through data sources); also see [column 39, lines 19-34] for a request to blend datasets); “examining a plurality of candidate datasets in dependence on metadata indices describing associations among columns within the candidate datasets” (see Oberbreckling et al., [column 29, lines 38-45] for processing/examining data sets using profile data (i.e., metadata indices); also see [column 40, lines 1-23]; and “merging the candidate datasets into the production dataset in a manner that varies according to the received user objective specification” (see Oberbreckling et al., [column 8, line 62 to column 9, line 15], [column 27, lines 10-19] and [column 27, line 65 to column 28, line 5] for merging data sets based on recommendation(s) selected by a user). In case that Oberbreckling et al. does not explicitly teach receiving a user objective specification of a production dataset as equivalently recited. On the other hand, Dugan et al. explicitly teaches receiving a user objective specification of a production dataset (see Dugan et al., [column 8, lines 27-45] for receiving transformation code from a user, wherein the transformation code specifying one or more target columns of the target dataset(s) can be interpreted as user objective specification of the target/production dataset as recited). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Duran et al.'s teaching to Oberbreckling et al.’s system by implementing a feature for receiving a user objective specification of a production dataset. Ordinarily skilled artisan would have been motivated to do so to provide Oberbreckling et al.’s system with an effective way to specify the production dataset to be generated. In addition, both of the references (Oberbreckling et al. and Dugan et al.) teach features that are directed to analogous art and they are directed to the same field of endeavor, such as, processing datasets to generate data/metadata associated with datasets and forming/generating a dataset by merging/joining datasets. This close relation between both of the references highly suggests an expectation of success when combined. As to claim 9, this claim is rejected based on the same arguments as above to reject claim 8 and is similarly rejected including the following: Oberbreckling et al. as modified by Dugan et al. teaches: “wherein the user objective specification includes at least one constraint selected from: a required set of columns, a semantic tag, a column value range, and a sample count” (see Oberbreckling et al., [column 11, lines 40-46] for receiving a data enrichment request from the client (i.e., a user), the data enrichment request (i.e., user objective specification) identifies a data source and/or particular data (e.g., tables, columns or other data available through data sources) for creating/obtaining a dataset; also see Dugan et al., [column 8, lines 27-45] for receiving transformation code from a user, wherein the transformation code specifying one or more target columns of the target dataset(s) can be interpreted as user objective specification of the target/production dataset as recited). As to claim 10, this claim is rejected based on the same arguments as above to reject claim 8 and is similarly rejected including the following: Oberbreckling et al. as modified by Dugan et al. teaches: “presenting a ranked list of candidate dataset groups satisfying the user objective specification on a user interface for user selection” (see Oberbreckling et al., [column 30, line 66 to column 31, line 27] for ranking and displaying ranked column pairs for selection by a user to merge datasets, wherein each column pair representing a suggestion for a possible join of the datasets can be interpreted as equivalent to a candidate dataset group as recited). As to claim 11, this claim is rejected based on the same arguments as above to reject claim 8 and is similarly rejected including the following: Oberbreckling et al. as modified by Dugan et al. teaches: “wherein the merging comprises selecting a merge strategy from among a plurality of available merge strategies, the merge strategy being selected in accordance with the user objective specification” (see Oberbreckling et al., [column 33, lines 8-24] for displaying different types of join functions (i.e., merge strategies) for selection; also see Dugan et al., [column 8, lines 27-45] for transformation code (i.e., user objective specification) specifying how to transforming/merging data from source datasets into target datasets). As to claim 12, this claim is rejected based on the same arguments as above to reject claim 8 and is similarly rejected including the following: Oberbreckling et al. as modified by Dugan et al. teaches: “wherein examining the candidate datasets comprises performing semantic similarity matching between terms in the user objective specification and column names of the candidate datasets” (see Oberbreckling et al., [column 23, lines 5-17] for performing a semantic similarity between datasets; also see [column 22, lines 11-20]; and also see Dugan et al., [column 8, lines 27-45] and [column 20, lines 20-26]). As to claim 13, this claim is rejected based on the same arguments as above to reject claim 8 and is similarly rejected including the following: Oberbreckling et al. as modified by Dugan et al. teaches: “wherein examining the candidate datasets comprises performing semantic similarity matching between terms in the user objective specification and column names of the candidate datasets, and wherein the semantic similarity matching is performed using at least one of: word embedding vector distances, N-gram comparisons, and abbreviation expansion” (see Oberbreckling et al., [column 23, lines 5-17] for performing a semantic similarity between datasets; also see [column 22, lines 11-20 and lines 29-38] for a similarity metric based on cosine similarity or distance using Word2Vec; and also see Dugan et al., [column 8, lines 27-45] and [column 20, lines 20-26]). As to claim 14, this claim is rejected based on the same arguments as above to reject claim 8 and is similarly rejected including the following: Oberbreckling et al. as modified by Dugan et al. teaches: “automatically merging the candidate datasets without user confirmation in response to receiving a predefined user objective specification” (see Oberbreckling et al., [column 27, lines 14-19] for merging two datasets using one or more transform scripts). As to claim 15, Oberbreckling et al. teaches: “A computer implemented method” (see Oberbreckling et al., Abstract) comprising: “generating a plurality of versioned production datasets, each version produced by merging different subsets of columns from a plurality of source datasets” (see Oberbreckling et al., [column 32, lines 28-30] for generating datasets by merging the datasets based on column pair relationship); “computing a respective composite dataset score for each versioned production dataset in dependence on metadata indices that quantify associations between the merged columns” (see Oberbreckling et al., [column 30, line 65 to column 31, line 27] for computing a score for each column pair, each column pair represents a suggestion for a possible join for the datasets to generate a dataset (i.e., a versioned production dataset)); “presenting the plurality of versioned production datasets and their respective composite dataset scores for display in a user interface” (see Oberbreckling et al., [column 30, line 65 to column 31, line 27] for displaying a set of column pairs (i.e., representing a version of a result dataset) and their corresponding score); and “initiating an automated selection, deployment, or downstream processing of a versioned production dataset based on a user choice or predefined selection rule” (see Oberbreckling et al., [column 7, lines 25-44] for providing data source enrichments (i.e., generated datasets) to other systems/services for processing). In case, Oberbreckling et al. does not explicitly teach a plurality of versioned datasets. On the other hand, Dugan et al. teaches a plurality of versioned datasets (see Dugan et al., [column 7, lines 28-30] for a datastore has versioned datasets). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Duran et al.'s teaching to Oberbreckling et al.’s system by implementing a plurality of versioned production datasets. Ordinarily skilled artisan would have been motivated to do so to provide Oberbreckling et al.’s system with an effective way to modify how a production dataset is generated from merging datasets. In addition, both of the references (Oberbreckling et al. and Dugan et al.) teach features that are directed to analogous art and they are directed to the same field of endeavor, such as, processing datasets to generate data/metadata associated with datasets and forming/generating a dataset by merging/joining datasets. This close relation between both of the references highly suggests an expectation of success when combined. As to claim 16, this claim is rejected based on the same arguments as above to reject claim 15 and is similarly rejected including the following: Oberbreckling et al. as modified by Dugan et al. teaches: “wherein computing the composite dataset score comprises applying weighted factors including at least one of: a common column factor, an association strength factor, and a semantic similarity factor” (see Oberbreckling et al., [column 29, lines 53-65]). As to claim 17, this claim is rejected based on the same arguments as above to reject claim 15 and is similarly rejected including the following: Oberbreckling et al. as modified by Dugan et al. teaches: “wherein presenting the versioned production datasets comprises ranking the versioned production datasets in accordance with the composite dataset scores” (see Oberbreckling et al., [column 30, line 66 to column 31, line 9] for displaying a ranked set of column pairs based on a score associated with each column pair). As to claim 18, this claim is rejected based on the same arguments as above to reject claim 15 and is similarly rejected including the following: Oberbreckling et al. as modified by Dugan et al. teaches: “wherein initiating the automated selection comprises deploying the selected versioned production dataset for training a machine learning model” (see Oberbreckling et al., [column 31, lines 60-65] for using data in machine learning (e.g., training)). As to claim 19, this claim is rejected based on the same arguments as above to reject claim 15 and is similarly rejected including the following: Oberbreckling et al. as modified by Dugan et al. teaches: “wherein initiating the automated selection comprises deploying the selected versioned production dataset for testing an application programming interface (API)” (see Oberbreckling et al., [column 7, lines 25-44] for providing data source enrichments (i.e., generated datasets) to other systems/services for processing (e.g., testing, training, etc.)). As to claim 20, this claim is rejected based on the same arguments as above to reject claim 15 and is similarly rejected including the following: Oberbreckling et al. as modified by Dugan et al. teaches: “wherein initiating the automated selection comprises controlling an enterprise process in dependence on predictions generated by a machine learning model trained using the selected versioned production dataset” (see Oberbreckling et al., [column 7, lines 25-44] for providing data source enrichments (i.e., generated datasets) to other systems/services for processing (e.g., analyzing, testing, training, etc.)). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHUONG THAO CAO whose telephone number is (571)272-2735. The examiner can normally be reached Monday - Friday: 9:00 am - 6:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amy Ng can be reached at 571-270-1698. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Phuong Thao Cao/Primary Examiner, Art Unit 2164
Read full office action

Prosecution Timeline

Sep 30, 2025
Application Filed
Aug 06, 2026
Non-Final Rejection mailed — §101, §103, §112
Sep 16, 2026
Applicant Interview (Telephonic)
Sep 16, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748762
TECHNIQUES FOR AUTOMATICALLY INFERRING INTENTS OF SEARCH QUERIES
3y 5m to grant Granted Sep 29, 2026
Patent 12737332
DATA SAMPLING METHOD THAT MAINTAINS ACCURACY FOR DATA ANALYSIS
3y 9m to grant Granted Sep 15, 2026
Patent 12730788
SYSTEMS AND METHODS FOR INTERLEAVING SEARCH RESULTS
2y 5m to grant Granted Sep 08, 2026
Patent 12682300
FOOD DATA ACCESS AND DELIVERY SYSTEM
5y 6m to grant Granted Jul 14, 2026
Patent 12657556
ELECTRONIC RECEIPT MANAGER APPARATUSES, METHODS AND SYSTEMS
2y 11m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
93%
With Interview (+14.6%)
2y 11m (~1y 11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 778 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month