Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a mental process without significantly more.
Representative claim 1 recites:
“A data lineage optimizing system comprising:
a processor; and
a memory operatively coupled with the processor, the processor being configured to:
remove, via a simplification process, a redundant data operation from a data lineage graph of a dataset in a data lake;
identify, via an alteration process, an alternate transformation as a candidate to replace an original transformation in the data lineage graph,
calculate a first value of a quality metric associated with the alternate transformation, wherein the alternate transformation causes the processor to transform an original dataset into a first resulting dataset; and
calculate a second value of the quality metric associated with the original transformation, wherein the original transformation causes the processor to transform the original dataset into a second resulting dataset, and wherein the quality metric comprises one of a reliability, a consistency, or an accuracy of a transformation;
in response to the first value of the quality metric associated with the alternate transformation exceeding the second value of the quality metric associated with the original transformation, replace the original transformation with the alternate transformation in the data lineage graph;
modify, in response to completion of the simplification process and the alteration process, at least one of an Extract, Transform, Load (ETL) process or an Extract, Load Transform (ELT) process configured to process data retrieved from the data lake; and
iteratively apply the simplification process and the alteration process to at least one additional datasets of the data lake.”
Independent Claims 15 and 18 recite similar subject matter.
The claim is directed to a mental process because the “simplification process,” “alteration process,” calculation of quality metrics, replacement of an original transformation with an alternate transformation, modification of an ETL or ELT process, and applying the simplification process and the alteration process to at least one additional dataset of the data lake all appear to be data analysis and data judgment processes that could be done by a human being with a generic computer. It is noted that, while one transformation may be replaced by another transformation in a data lineage graph, there is no functional result to this. While at least one of an ETL or ELT process is modified, it is noted that no data lineage graph is never claimed to actually transform data or be used in the data migration system.
A human being, equipped with pen and paper or a generic computer, is capable of performing data analyses and data judgment to modify a data lineage graph.
The additional elements in the claim include a processor and memory (claim 1), processor (claim 15), and computer-readable medium (claim 18).
This judicial exception is not integrated into a practical application because the claimed additional elements do not appear to improve the processing of a computer, require the use of a specific machine, effect a transformation or reduction of a particular article to a different state or thing, or provide a technological solution to a technological problem.
The processors, memory, and computer-readable medium are recited at a high level of generality. They appear to be generic computing elements or hardware elements. The recitation of generic hardware or generic computing elements is little more than using a computer to perform an abstract idea, see MPEP 2106.05(f)(2).
It is noted that none of the additional elements appear to improve the processing of a computer, require the use of a specific machine, effect a transformation or reduction of a particular article to a different state or thing, or provide a technological solution to a technological problem. As such, none of the additional elements appear to integrate the judicial exception into a practical application.
None of the additional elements are sufficient to amount to significantly more than the judicial exception, in part or in whole.
The recitation of generic hardware or computing elements of the processor, memory, and computer-readable medium are little more than using a computer to perform an abstract idea, see MPEP 2106.05(f)(2).
None of the additional elements, in part or in whole, appear to improve the processing of a computer, require the use of a particular machine, effect a transformation or reduction of a particular article to a different state or thing, or add a specific limitation other than what is well understood, routine, or conventional. As such, none of the additional elements appears to be, in part or in whole, significantly more than the judicial exception.
Dependent claims 2-14, 16-17, and 19-20 are merely directed towards additional limitations that further define data types or further describe analyses that will occur. It is noted that the claimed data definitions and data analysis and extraction steps do not appear to include additional elements that incorporate the claimed subject matter into a practical application. The dependent claims also do not include additional elements that, in part or in whole, appear to be significantly more than the abstract idea.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claims 1-2 and 15-18 are rejected under 35 U.S.C. 103 as being unpatentable over Dickie et al. (US Pre-Grant Publication 2025/0181319) in view of Fan et al. (US Pre-Grant Publication 2021/0303585).
As to claim 1, Dickie teaches a data lineage optimizing system comprising
a processor (see paragraph [0085]); and
a memory operatively coupled with the processor (see paragraph [0085]), the processor being configured to:
remove, via a simplification process, a redundant data operation from a data lineage graph of a dataset in a data lake (see paragraph [0063]-[0064] for a definition of a dataflow graph describing operations, wherein components are nodes that accept input fields and produce output fields. Paragraph [0097]-[0098] discuss ways to transform a dataflow graph to be more efficient. Among the methods include removing redundant components, see paragraph [0098]. As noted in paragraph [0101], the storage in Dickie may comprise multiple databases in a data lake);
identify, via an alteration process, an alternate transformation as a candidate to replace an original transformation in the data lineage graph (see paragraphs [0097]-[0098] for identifying alternate data processing operations),
…
wherein the alternate transformation causes the processor to transform an original dataset into a first resulting dataset (see paragraphs [0097]-[0098] and [0100]. Both the original and optimized dataflow graph transform an original dataset into a resulting dataset)
…
wherein the original transformation causes the processor to transform the original dataset into a second resulting dataset… (see paragraphs [0097]-[0098] and [0100]. Both the original and optimized dataflow graph transform an original dataset into a resulting dataset)
in response to … the alternate transformation … [optimizing] the original transformation, replace the original transformation with the alternate transformation in the data lineage graph (see paragraphs [0097]-[0098]. The original transformation undergoes an optimization process in which the alternate, replacement transformation is an optimized version of the original transformation).
Dickie does not teach to:
calculate a first value of a quality metric associated with the alternate transformation; and
calculate a second value of the quality metric associated with the original transformation, and wherein the quality metric comprises one of a reliability, a consistency, or an accuracy of a transformation; and
in response to the first value of the quality metric associated with the alternate transformation exceeding the second value of the quality metric associated with the original transformation, replace the original transformation with the alternate transformation in the data lineage graph.
modify, in response to the completion of the simplification process and the alteration process, at least one of an Extract, Transform, Load (ETL) process or an Extract, Load, Transform (ELT) process configured to process data retrieved from the data lake; and
iteratively apply the simplification process and the alteration process to at least one additional datasets of the data lake.
Fan teaches:
calculate a first value of a quality metric associated with the alternate transformation, wherein the alternate transformation transforms an original dataset into a first resulting dataset (see Fan paragraph [0071]-[0073]. Fan teaches to analyze a data pipeline. Fan also teaches, as shown in [0073], to consider modifications to the data pipeline plan. Metrics between an original plan and a modified plan may be compared to determine whether a reduction or an increase in a metric occurs as a result of the modification); and
calculate a second value of the quality metric associated with the original transformation, wherein the original transformation transforms the original dataset into a second resulting dataset, and wherein the quality metric comprises one of a reliability, a consistency, or an accuracy of a transformation (see Fan paragraph [0073]. As noted above, metrics between an unmodified combination of datasets and a modified combination of datasets may be compared. The metrics include reliability of a transformation, notably, an increase in a delivery speed. An increased delivery speed is more reliable than a decreased or slower delivery speed. Also see paragraph [0017], which indicates that metrics may be gathered for a pipeline, including freshness, a measure of reliability and consistency, and quality); and
in response to the first value of the quality metric associated with the alternate transformation exceeding the second value of the quality metric associated with the original transformation, replace the original transformation with the alternate transformation in the data lineage graph (see Fan paragraph [0073]. A modification that improves upon a metric will be selected and implemented).
modify, in response to the completion of the simplification process and the alteration process, at least one of an Extract, Transform, Load (ETL) process or an Extract, Load, Transform (ELT) process configured to process data retrieved from the data lake (see Fan paragraph [0002]. The data pipelines of Fan may be used to define ETL systems. As noted in paragraph [0101] of Dickie, the input storage may comprise multiple databases in a data lake); and
iteratively apply the simplification process and the alteration process to at least one additional datasets of the data lake (see Fan paragraph [0077] and [0126]. The plan may be applied to multiple pipelines to establish multiple pipelines in accordance with the plan).
It would have been obvious to one of ordinary skill in the art before the earliest filing date of the invention to have modified Dickie by the teachings of Fan because both references are directed towards modifying and improving data analysis. Fan provides to Dickie additional monitoring steps to ensure that proposed updates to data analysis are more efficient and accurate than existing models, which will improve the ability of Dickie to manage and optimize multiple dataflows.
As to claim 2, Dickie as modified by Fan teaches the data lineage optimizing system of claim 1, the processor is further configured to:
execute the simplification process and the alteration process in parallel on additional datasets of the data lake (see Fan paragraph [0073]).
As to claims 15 and 18, see the rejection of claim 1.
As to claim 16, Dickie teaches the method of claim 15, wherein iteratively executing the simplification process and the alteration process further comprises:
serially executing by the processor, the simplification process, and the alteration process on additional datasets of the data lake (see Dickie paragraphs [0094]-[0095] and [0097]-[0098]. Dataflow graphs may be generated and processed in response to user commands).
As to claim 17, Dickie teaches method of claim 16, wherein serially executing the simplification process, and the alteration process further comprises:
replacing, by the processor, the original transformation with the alternate transformation in the data lineage graph of the first resulting dataset via the execution of the alteration process (see Dickie paragraphs [0097]-[0098]).
Claims 3-4 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Dickie et al. (US Pre-Grant Publication 2025/0181319) in view of Fan et al. (US Pre-Grant Publication 2021/0303585), and further in view of Abdul-Jawad et al. (US Patent 10,775,976).
As to claim 3, Dickie as modified teaches the data lineage optimizing system of claim 1.
Dickie as modified does not teach wherein the processor is further configured to:
generate new abstract syntax trees (ASTs) from the simplification process and the alteration processes.
Abdul-Jawad teaches wherein the processor is further configured to:
generate new abstract syntax trees (ASTs) from the simplification process and the alteration processes (see Abdul-Jawad 144:25-50).
It would have been obvious to one of ordinary skill in the art before the earliest filing date of the invention to have modified Dickie by the teachings of Abdul-Jawad because both references are directed towards improving dataflows. Abdul-Jawad provides to Dickie an additional way of representing data flows to understand operations that occur.
As to claim 4, Dickie as modified by Abdul-Jawad teaches the data lineage optimizing system of claim 3, wherein the processor is further configured to:
translate the new abstract syntax trees (ASTs) into executable code that modifies the at least one of the Extract, Transform, Load (ETL) or Extract, Load, Transform (ELT) processes for the data lake (see Abdul-Jawad 146:31-61. Abdul-Jawad teaches to send the ASTs to the intake system, which performs data processing. See Abdul-Jawad 8:11-18 for a data intake system that extracts data).
As to claim 19, Dickie as modified teaches the non-transitory computer-readable medium of claim 18, wherein the instructions further cause the processor to:
Serially execute the simplification process and the alteration process on additional datasets of a data lake (see Dickie paragraphs [0094]-[0095] and [0097]-[0098] and the rejection of claim 16); and
Dickie does not teach to generate new abstract syntax trees (ASTs) from the serially executed simplification process and alteration processes.
.Abdul-Jawad teaches to generate new abstract syntax trees (ASTs) from the serially executed simplification process and alteration processes (see Abdul-Jawad 144:25-50).
It would have been obvious to one of ordinary skill in the art before the earliest filing date of the invention to have modified Dickie by the teachings of Abdul-Jawad because both references are directed towards improving dataflows. Abdul-Jawad provides to Dickie an additional way of representing data flows to understand operations that occur.
As to claim 20, Dickie teaches the non-transitory computer-readable medium of claim 19, wherein the instructions further cause the processor to:
translate the new abstract syntax trees (ASTs) into executable code that modifies the at least one of the Extract, Transform, Load (ETL) or Extract, Load, Transform (ELT) processes of the data lake (see Abdul-Jawad 146:31-61. Abdul-Jawad teaches to send the ASTs to the intake system, which performs data processing. See Abdul-Jawad 8:11-18 for a data intake system that extracts data).
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Dickie et al. (US Pre-Grant Publication 2025/0181319) in view of Fan et al. (US Pre-Grant Publication 2021/0303585), and further in view of Venkataramani et al. (US Patent 10,114,917)
As to claim 5, Dickie as modified teaches the data lineage optimizing system of claim 1.
Dickie does not clearly teach wherein the processor is further configured to:
calculate a transform complexity of a transformation used to generate at least one column of the dataset; and
extract dependencies of the at least one column of the dataset from an abstract syntax tree (AST) of the transformation.
Venkataramani teaches:
calculate a transform complexity of a transformation used to generate at least one column of the dataset (see 31:41-65. A data model, including a data flow, may be analyzed. Characteristics of each block of the data model may be analyzed, including with a view towards data complexity. It is noted that Dickie, cited above, teaches wherein data fields represent columns, see paragraph [0005], and a data flow being a data flow graph, [0097]-[0098]); and
extract dependencies of the at least one column of the dataset from an abstract syntax tree (AST) of the transformation (see 21:62-22:6 for creating an abstract data flow for an intermediate representation of a data model. It is noted that Dickie, cited above, teaches wherein data fields represent columns, see paragraph [0005]).
It would have been obvious to one of ordinary skill in the art before the earliest filing date of the invention to have modified Dickie by the teachings of Venkataramani because both references are directed towards improving dataflows. Venkataramani provides to users Dickie an additional way of analyzing data flows to better understand a data lineage and operations that occur.
Claims 6-7 are rejected under 35 U.S.C. 103 as being unpatentable over Dickie et al. (US Pre-Grant Publication 2025/0181319) in view of Fan et al. (US Pre-Grant Publication 2021/0303585), in view of Venkataramani et al. (US Patent 10,114,917), and further in view of Dickie et al. (US Pre-Grant Publication 2019/0370407, hereinafter “Dickie ‘407”).
As to claim 6, Dickie as modified teaches the data lineage optimizing system of claim 5.
Dickie does not teach wherein the processor is further configured to:
continue calculating the transform complexity and extracting of dependencies for a plurality of columns including the at least one column of the dataset until a sum of the transform complexities of the plurality of columns exceeds a transform complexity threshold.
Dickie ‘407 teaches wherein the processor is further configured to:
continue calculating the transform complexity and extracting of dependencies for a plurality of columns including the at least one column of the dataset until a sum of the transform complexities of the plurality of columns exceeds a transform complexity threshold (see Dickie ‘407 paragraph [0093]. A series of nodes in a data flow may be judged for “complexity,” such as whether there is a series of nodes that could be combined. There is minimum threshold length).
It would have been obvious to one of ordinary skill in the art before the earliest filing date of the invention to have modified Dickie by the teachings of Dickie ‘407 because both references are directed towards improving dataflows. Dickie ‘407 provides to Dickie an additional way of improving data flows by combining elements making a data flow more efficient.
As to claim 7, Dickie as modified by Dickie ‘407 teaches the data lineage optimizing system of claim 6, the processor is further configured to:
identify one of a plurality of datasets of the data lake processed in a step immediately preceding the sum of the transform complexities exceeding the transform complexity threshold as a source dataset directly receiving a dependency of the dataset (see Dickie ‘407 paragraph [0093]).
Claims 8-9 are rejected under 35 U.S.C. 103 as being unpatentable over Dickie et al. (US Pre-Grant Publication 2025/0181319) in view of Fan et al. (US Pre-Grant Publication 2021/0303585), in view of Polleri et al. (US Pre-Grant Publication 2024/0320303).
As to claim 8, Dickie teaches the data lineage optimizing system of claim 1.
Dickie does not teach wherein to execute the alteration process that identifies the alternate transformation, the processor is configured to:
extract quality metrics of data generated by the original transformation and data generated by the alternate transformation.
Polleri teaches wherein to execute the alteration process that identifies the alternate transformation, the processor is configured to:
extract quality metrics of data generated by the original transformation and data generated by the alternate transformation (see paragraph [0033]. Polleri monitors pipelines or data workflows for compliance with quality of service dimensions, see abstracts. Note that this may include a comparison of multiple workflows. Also see paragraphs [0067]-[0068] for a selection of multiple pipelines to compare and order).
It would have been obvious to one of ordinary skill in the art before the earliest filing date of the invention to have modified Dickie by the teachings of Polleri because both references are directed towards managing data flows. Polleri simply adds to Dickie the ability to compare two data flows and sort them based on how well they match a quality requirement. This will result in better data flows being identified.
As to claim 9, Dickie as modified by Polleri teaches the data lineage optimizing system of claim 8, wherein the quality metric further comprises at least one of: quality checks performed, a service level agreement guaranteed, or a frequency of fulfillment of the service level agreement (see Polleri paragraphs [0067]-[0070]).
Claims 10-11 and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Dickie et al. (US Pre-Grant Publication 2025/0181319) in view of Fan et al. (US Pre-Grant Publication 2021/0303585), in view of Polleri et al. (US Pre-Grant Publication 2024/0320303), and further in view of Shapur et al. (US Pre-Grant Publication 2020/0081899).
As to claim 10, Dickie as modified by Polleri teaches the data lineage optimizing system of claim 8.
Dickie does not teach wherein to identify the alternate transformation, the processor is configured to:
generate statistical column metrics vectors for the first resulting dataset, and other datasets including the original dataset of the data lake; and
calculate similarities between the statistical column metrics vector of the first resulting dataset and the statistical column metrics vectors of other datasets including the original dataset of the data lake.
Shapur teaches:
wherein to identify the alternate transformation, the processor is configured to:
generate statistical column metrics vectors for the first resulting dataset, and other datasets including the original dataset of the data lake (see paragraphs [0072]-[0075]. Shapur generates vectors for source and target columns and compares them); and
calculate similarities between the statistical column metrics vector of the first resulting dataset and the statistical column metrics vectors of other datasets including the original dataset of the data lake (see paragraphs [0072]-[0075]. Shapur determines similarities from the source and target vectors).
It would have been obvious to one of ordinary skill in the art before the earliest filing date of the invention to have modified Dickie by the teachings of Shapur because both references are directed towards transforming data. Shapur simply adds to Dickie the ability to compare identify additional data metrics that will help to determine how well a source column maps to a target column. This will result in better data flows being identified.
As to claim 11, Dickie as modified by Shapur teaches the data lineage optimizing system of claim 10, wherein to calculate the similarities, the processor is configured to:
calculate the similarities with one of Euclidean distance or cosine similarity measures (see Shapur paragraph [0074]).
As to claim 13, Dickie as modified by Polleri teaches the data lineage optimizing system of claim 10, wherein to identify the alternate transformation, the processor is configured to:
select a transformation that maximizes the similarities and improves on the quality metrics of the original transformation as the alternate transformation (see Polleri paragraphs [0067]-[0070]).
As to claim 14, Dickie as modified by Fan teaches the data lineage optimizing system of claim 13, wherein the alternate transformation removes an existing dataset from a data lineage of an original dataset and further adds one or more of a new dataset and a new operation to a data lineage of the original dataset (see Fan paragraph [0073] for adding a dataset and operation. Also see Dickie paragraph [0097]-[0098] for removing a dataset).
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Dickie et al. (US Pre-Grant Publication 2025/0181319) in view of Fan et al. (US Pre-Grant Publication 2021/0303585), in view of Polleri et al. (US Pre-Grant Publication 2024/0320303), in view of Shapur et al. (US Pre-Grant Publication 2020/0081899), and further in view of Venkataramani et al. (US Patent 10,114,917).
As to claim 12, Dickie as modified teaches the data lineage optimizing system of claim 10.
Dickie does not teach wherein to identify the alternate transformation, the processor is configured to:
identify based at least on abstract syntax trees (ASTs) of the original transformation, column dependencies of the first resulting dataset.
Venkataramani teaches wherein to identify the alternate transformation, the processor is configured to:
identify based at least on abstract syntax trees (ASTs) of the original transformation, column dependencies of the first resulting dataset (see 21:62-22:6 for creating an abstract data flow for an intermediate representation of a data model, including dependencies. It is noted that Dickie, cited above, teaches wherein data fields represent columns, see paragraph [0005]).
It would have been obvious to one of ordinary skill in the art before the earliest filing date of the invention to have modified Dickie by the teachings of Venkataramani because both references are directed towards improving dataflows. Venkataramani provides to users Dickie an additional way of analyzing data flows to better understand a data lineage and operations that occur.
Response to Arguments
Applicant's arguments filed 26 May 2026 have been fully considered but they are not persuasive.
Response to Rejections under 35 USC 101
Applicant disagrees with the rejection under 35 USC 101 of the claimed invention being directed to an abstract idea of a mental process. Applicant then lists out each step of the claims.
Applicant argues that “that the claims are not directed to a mental process, but instead recite specific technological operations for automatically modifying executable ETL/ELT processing pipelines within a data lake environment based on iterative lineage optimization and transformation-quality evaluation.”
In response to this argument, it is noted merely automating an otherwise manual process is not indicative of an improvement to the functioning of a computer. See MPEP 2106.05(a)(I), the second item under the list of “Examples that the courts have indicated may not be sufficient to show an improvement in computer-functionality,”
iii. Mere automation of manual processes, such as using a generic computer to process an application for financing a purchase, Credit Acceptance Corp. v. Westlake Services, 859 F.3d 1044, 1055, 123 USPQ2d 1100, 1108-09 (Fed. Cir. 2017) or speeding up a loan-application process by enabling borrowers to avoid physically going to or calling each lender and filling out a loan application, LendingTree, LLC v. Zillow, Inc., 656 Fed. App'x 991, 996-97 (Fed. Cir. 2016) (non-precedential);
The claimed identification, calculation, replacement, modification, and application steps are data analysis or data judgment tasks that may all be performed by a human being equipped with a generic computer. Simply automating a manual series of analysis and judgments steps performed by a computer or that could be performed by a user is not is not sufficient to show a practical application.
Applicant argues that “In particular, the claims recite automated simplification and alteration processes that operate on machine-scale data lineage graphs, evaluate competing transformations using quality metrics, replace transformations within lineage structures, generate executable modifications to ETL/ELT processing logic, and iteratively optimize additional datasets within a data lake environment. Such operations are fundamentally rooted in computer technology and are not practically performable in the human mind.”
As noted in MPEP 2106.04(a)(2) III C, “claims can recite a mental process even if they are claimed as being performed on a computer.
The Supreme Court recognized this in Benson, determining that a mathematical algorithm for converting binary coded decimal to pure binary within a computer’s shift register was an abstract idea. The Court concluded that the algorithm could be performed purely mentally even though the claimed procedures "can be carried out in existing computers long in use, no new machinery being necessary." 409 U.S at 67, 175 USPQ at 675. See also Mortgage Grader, 811 F.3d at 1324, 117 USPQ2d at 1699 (concluding that concept of "anonymous loan shopping" recited in a computer system claim is an abstract idea because it could be "performed by humans without a computer").”
MPEP 2106.04(a)(2) III C 1-3 further elaborate on the idea that a claim may still be directed towards an abstract idea despite the use of a generic machine. Thus, though the claims may not be performed solely in a human mind, simplification and alteration processes, identification, calculation, replacement, modification, and iterative application steps may be performed by a human with a generic computer.
The fact that a problem exists in a computer environment is, on its own, not sufficient to declare an invention not directed towards a mental process.
Applicant argues that “For example, the claims recite modification of ETL/ELT processes configured to process data retrieved from a data lake, as well as iterative application of simplification and alteration processes across additional datasets. The claims further recite generation of new abstract syntax trees (ASTs) and translation of those ASTs into executable code that modifies ETL/ELT processes. These operations require automated manipulation of executable processing logic and distributed data-processing infrastructure, and therefore cannot practically be performed mentally.”
Applicant then cites paragraphs [0018] and [0036] of the Specification.
Applicant argues that “Applicant respectfully submits that the claims do not merely analyze or evaluate data lineage information. Rather, the claims recite specific mechanisms for automatically generating executable modifications to operational ETL/ELT infrastructure within a data lake environment. In particular, the claims recite automated restructuring of executable data-processing pipelines based on iterative lineage optimization and comparative transformation-quality evaluation across datasets.”
Examiner notes that abstract syntax trees are a representation of data in a particular format. A human being equipped with a generic computer or pen and paper is capable of such a conversion. Simply generating abstract syntax trees in response to a data manipulation operation is a data analysis process, and thus a mental process.
As noted above, neither mere “automation” of “manipulation of executable processing logic,” “restructuring of executable data-processing pipelines,” nor the execution of steps using generic computing components and computing environments is sufficient to remove a claim from consideration as an abstract idea.
Applicant argues that “The Examiner has alleged that the data lineage graph is never claimed to actually transform data or be used in the data migration system. However, the amended claims now expressly recite "modify, in response to completion of the simplification process and the alteration process, at least one of an Extract, Transform, Load (ETL) process or an Extract, Load, Transform (ELT) process configured to process data retrieved from the data lake." This limitation directly addresses the Examiner's concern by requiring operational modification of executable ETL/ELT processes that process data retrieved from the data lake.”
In response to this argument, an ELT or ETL process is being modified in this step. However, neither the ETL or ELT process is being executed or used after modification. It is merely “configured to process data,” but does not actually process data. According to the claim language, it is merely a data process flow design that has yet to be implemented. Therefore, no improvements occur in an ETL or ELT process.
To put this another way, while the ELT or ETL process may be designed to be more efficient by optimizing a data flow graph on paper, any improvement to the design of a data flow graph is not reflected in a computing system until the improved design is used.
Applicant argues that “The evaluation of a practical application requires consideration of whether the claims improve computer functionality or another technology or technical field. Applicant respectfully submits that the amended claims improve operation of distributed data-processing infrastructure by automatically restructuring executable ETL/ELT processing pipelines within a data lake environment through iterative lineage optimization and transformation-quality evaluation.”
Applicant continues, arguing that “The Specification explains that complex data pipelines and data lineage structures can introduce redundant transformations, inefficient dependencies, excessive intermediate processing operations, and degraded data quality. The claims address these technological problems through automated simplification and alteration processes that remove redundant operations, compare competing transformations using quality metrics, regenerate executable processing logic, and iteratively optimize additional datasets within the data lake environment.”
Applicant then cites paragraphs [0012], [0021], and [0036] of the specification.
As noted above, the “simplification and alteration processes that remove redundant operations,” as claimed, appears to be involve data analysis and data judgment steps. The fact that this process is “automated” does not integrate the data analysis and judgment process into a practical application or render it significantly more than the abstract idea.
As noted in the previous paragraph, while an ETL or ELT process may be modified as a result of the data analysis, the ETL and ELT processes are claimed as only “configured to process data” and are not actually implemented.
Applicant argues that “Thus, the claims recite a specific technological mechanism for reducing redundant execution of transformation operations within distributed data-processing pipelines, reducing intermediate processing stages, improving execution efficiency of ETL/ELT workflows, and improving operational reliability of data-processing infrastructure within a data lake environment. In particular, the claims do not merely evaluate data lineage information or recommend modifications. Rather, the claims automatically generate executable modifications to ETL/ELT processing logic through AST generation and executable code translation mechanisms. Accordingly, the claims recite concrete technological implementation of lineage-driven pipeline restructuring rather than mere information analysis. As noted in MPEP §2106.05(a), and as explained in Enfish, claims directed to "a specific implementation of a solution to a problem in the software arts" may demonstrate patent eligibility. Similar to Enfish, the present claims recite a specific improvement to computer functionality itself, namely automated optimization and executable restructuring of ETL/ELT data-processing pipelines operating within a data lake architecture, rather than merely using computers as tools to perform abstract analysis. The claims are therefore directed to specific technological mechanisms for modifying executable data-processing infrastructure and improving operation of distributed data-processing systems. Accordingly, Applicant respectfully submits that the claims integrate any alleged abstract idea into a practical application and are not directed to patent-ineligible subject matter under 35 U.S.C. § 101. Accordingly, Applicant submits that the claims are directed to patentable subject matter and respectfully requests withdrawal of the rejections under 35 U.S.C. § 101.”
Fundamentally, the design process of an ETL/ELT data flow is an abstract idea that a human being may do with a pen and paper or a generic machine to perform the necessary calculations. This is a mental process. Improvements to the design process may result in an improved design, however an improved mental process remains a mental process, albeit improved.
As discussed above, merely automating an otherwise manual design process does not provide an improvement to computing technology.
Additionally, the claim language only modifies ETL and ELT processes that are “configured to process data,” but never actually process data (including the extracting, transformation, and loading steps) within the scope of the claim. Thus, no improvement to an ETL or ELT process itself occurs within the scope of the claims.
Because the alleged improvements are directed automating data analysis and design and because this data analysis and design is not integrated into a practical application or significantly more than the data analysis process itself, Applicant’s arguments are unpersuasive.
Response to Rejections under 35 USC 103
Applicant recites the claim language, summarizes Dickie, then argues that “Dickie does not teach or suggest "modify, in response to completion of the simplification process and the alteration process, at least one of an Extract, Transform, Load (ETL) process or an Extract, Load, Transform (ELT) process configured to process data retrieved from the data lake" as recited by amended claim 1.””
Applicant continues, arguing that “Dickie describes a dataflow graph optimization that operates on abstract graph representations to improve computational efficiency of executing those graphs, but does not teach modifying actual ETL or ELT processes configured to process data retrieved from a data lake. Dickie teaches that the compiler module compiles a dataflow graph "into an executable software application program that can be executed by the data processing system" and that "the stored software application program may then be executed by the data processing system at a subsequent time." Furthermore, Dickie does not teach iteratively applying a simplification process and the alteration process to at least one additional datasets of the data lake (e.g., to optimize additional data lineages of the data lake). Rather, Dickie describes iteratively applying optimization rules to portions of a single dataflow graph, which is not the same as iteratively applying both a simplification process and an alteration process to at least one additional datasets of a data lake to optimize multiple different data lineage graphs within a data lake. Therefore, Dickie fails to teach or suggest at least the features of claim 1 discussed above.”
In response to this argument, it is noted that Fan, as cited above, is relied upon to teach the amended limitations.
In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986).
Applicant argues that “Fan, however, does not teach or suggest "modify, in response to completion of the simplification process and the alteration process, at least one of an Extract, Transform, Load (ETL) process or an Extract, Load, Transform (ELT) process configured to process data retrieved from the data lake" and "iteratively apply the simplification process and the alteration process to at least one additional datasets of the data lake" of claim 1. Fan describes data delivery optimization including creating roadmaps for delivering data from sources to targets, reducing pipeline components, network bandwidth, and latency, but does not describe modifying ETL or ELT processes in response to completion of simplification and alteration processes that optimize data lineage graphs. In particular, Fan provides for creating delivery plans for configuring data pipeline components but does not mention modification of processes for extracting, translating, or loading data from a data lake. Furthermore, while Fan describes combining information models for multiple data requests, Fan does not describe iteratively applying both a simplification process that removes redundant data operations from a data lineage graph and an alteration process that identifies alternate transformations based on quality metrics to at least one additional datasets of a data lake. Therefore, applicant submits that Fan fails to cure the deficiencies of Dickie with respect to claim 1 and that the combination of Dickie and Fan fails to teach or suggest the claimed features because neither reference, alone or in combination, discloses: (1) modifying an ETL or ELT process in response to completion of a simplification process and an alteration process, wherein the ETL or ELT process is configured to process data retrieved from a data lake; or (2) iteratively applying both a simplification process and an alteration process to at least one additional datasets of a data lake.”
In response to this argument, it is noted that a combination of Fan and Dickie does consider altering ETL or ELT processes (see Fan paragraph [0002]. The data pipelines of Fan that are altered may be used to define ETL systems. As noted in paragraph [0101] of Dickie, the input storage may comprise multiple databases in a data lake).
Fan also considers iteratively applying modification processes to multiple pipelines (see Fan paragraph [0077] and [0126]. The plan may be applied to multiple pipelines to establish multiple pipelines in accordance with the plan).
Thus, the references combined do teach the claimed subject matter to the extent claimed.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHARLES D ADAMS whose telephone number is (571)272-3938. The examiner can normally be reached M-F, 9-5:30 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aleksandr Kerzhner can be reached at 5712701760. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CHARLES D ADAMS/Primary Examiner, Art Unit 2165