DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
The Office has withdrawn the previous formalities objection to claim 28 in light of the claim amendment.
Applicants’ arguments, filed 6/17/2026, concerning the previous rejection of the claims under 35 USC §§101 and 103 have been fully considered but they are not persuasive.
Regarding the previous rejection of the claims under 35 USC §101, Applicants first argue on pages 7-8 that newly amended language directed to producing code / programming that uses/accesses data should not be considered an abstract concept that can be performed in the human mind.
The Office respectfully disagrees. As recognized in the first paragraph of the remarks on page 8: The courts consider a mental process as being able to be “performed in the human mind, or by a human using a pen and paper”. A programming language [code] is merely text that may be written using a pen and paper. And, code normally is associated with data.
Regarding the previous rejection of the claims under 35 USC §101, Applicants further assert on pages 8-10 the newly amended language as being evidence of an improvement to the technology, and therefore integrated into a practical application.
The Office respectfully disagrees. Generic computing elements (hardware/software) performing generic computer functions to apply an abstract idea do not amount to significantly more than the abstract idea of organizing information through mathematical correlations. It is noted that the Internet/computer limitations are simply a field of use that attempt to limit the abstract idea to a particular technological environment and do not add significantly more than the abstract idea itself. Viewing the limitations as a combination does not add anything further than looking at the limitations individually.
Regarding the previous rejection of the claims under 35 USC §101, Applicants assert on page 10 that Examiners should not assert claim rejections under 35 USC §101 unless said Examiners believe that it is more likely than not that the claims run afoul of 35 USC §101.
The Office respectfully asserts that the claims merit a rejection under 35 USC §101 using a preponderance of the evidence standard.
Therefore, the previous rejection of the claims under 35 USC §101 is believed to be reasonable and thus maintained.
Applicants’ arguments, filed 6/17/2026, concerning the previous rejection of the claims under 35 USC §103(a) have been fully considered but they are not persuasive.
Regarding the previous rejections of independent claim 24 under 35 USC §103, Applicants argue on pages 10-11 that the references do not teach the use of a language model.
The Office respectfully disagrees, noting that the references as a whole teach the claimed subject matter. First, it is noted that in light of Applicants’ own specification at [0034] indicates that such a limitation is Applicant Admitted Prior Art (AAPA). Applicants’ disclosure at paragraph 0034 discusses the use of GPT-3 as a language model (i.e., a competitor’s prior art software). Additionally, it is noted that at least the Romero Calvo reference discusses the use of trained ML model for the conversion of a NLQ to a SQL query.
Also regarding the previous rejections of independent claim 24 under 35 USC §103, Applicants argue on pages 10-11 that the references do not teach the claim language indicating that the data pipeline is written in a programming language as associated /applicable to a data set because the pipeline contains operators.
The Office respectfully disagrees, noting that the references as a whole teach the claimed subject matter. It is noted that the operators are programming code, and that generally operators take data as their arguments/operands.
Therefore, it is believed that claim 24 was reasonably rejected under 35 USC 103.
Applicants further argue on page 11 that the independent claims reciting substantially similar limitations, and all dependent claims are allowable for the reasons argued above.
The Office respectfully disagrees, and counter-asserts the rationale set forth above.
Claim Rejections – 35 U.S.C. § 101
35 U.S.C. § 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 24-43 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to non-statutory subject matter.
At step 1: claim 24 is directed to a “method” and thus directed to a statutory category, claim 36 is directed to a “system” and thus directed to a statutory category. Claim 43 is directed to a “non-transitory computer readable storage medium” claim thus directed to a statutory category.
At step 2a prong 1: claims 24, 36 and 43 recite the limitation that is directed to an abstract idea, “generating the data pipeline in a programming language and applicable to a dataset based at least in part on the generated query in the standard query language, the data pipeline comprising one or more data pipeline elements, at least one data pipeline element of the one or more data pipeline elements corresponding to a query component of the generated query in the standard query language” as drafted recites a mentally perforable process as one can generate a data pipeline of elements corresponding to a query component mentally or with the aid of pencil and paper.
At step 2a prong 2: Claims 24, 36 and 43 recite the following additional elements, “receiving a natural language (NL) query …” and “receiving a model result based upon the NL query …”. Such language represents insignificant extra-solution activity as retrieval/receiving of data (i.e. mere data gathering) such as 'obtaining information' as identified in MPEP 2106.05(g) and does not provide integration into a practical application.
Additionally, claims 24, 36 and 43 recite the additional elements: “one or more processors” (claim 24), “one or more processors” and “one or more memories” (claim 36), and “one or more processors” and a “storage medium” (claim 43). These elements are high-level recitations of generic computer components and represent mere instructions to apply on a computer as in MPEP 2106.05(f), which does not provide integration into a practical application.
Viewing the additional limitations together and the claims as a whole, nothing provides integration into a practical application.
At step 2b: the conclusions for the additional elements representing mere implementation using a computer are carried over and do not provide significantly more.
The two "receiving" limitations are identified as insignificant extra-solution activity above, and when re-evaluated these elements are well-understood, routine, and conventional as evidenced by the court cases in MPEP 2106.05(d)(II), "i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); … OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network);" and thus remain insignificant extra-solution activity that does not provide significantly more.
Therefore, each of these claims as a whole does not change this conclusion, and the claims are ineligible.
Claims 25-35 and 37-42 depend upon claims 24 and 36, respectively, and do not correct the issues set forth above. These claims essentially further recite data and generic/abstract processing modifications that may be performed mentally or via the aid of paper and pencil. Therefore, these claims are likewise rejected.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 24-33 and 35-43 are rejected under 35 U.S.C. §103 as being unpatentable over Romero Calvo et al (US Patent No. 12,124,440, hereafter referred to as “Romero”) in view of Immanuel Haffner et al. (“Fast Compilation and Execution of SQL Queries with WebAssembly”, arXiv, Cornell University Archive, document no: arXiv:2104.15098v2 [cs.DB], 3 May 2021, downloaded from: https://arxiv.org/abs/2104.15098, pp. 1-14, hereafter referred to as “Haffner”).
Regarding independent claim 24: Romero teaches A method for generating a data pipeline, the method comprising: receiving a natural language (NL) query, the NL query including one or more constraints associated with a target dataset; (See Romero Abstract, col. 2 line 62 – col. 3 line 7, and Figures 1 and 2 teaching the reception of a NLQ [natural language query] sentence comprised of “parts” that are necessary in order to compose a final SQL query, and are ultimately replaced by “key arguments”. See also, col. 9 line 59 – col. 10 line 14 discussing the use of an NLP-SQLQ engine that provides a pipeline of different stages/processes) receiving a model result generated based on the NL query, the model result including a generated query in a standard query language, the model result being generated using one or more computing models including a language model; (See Romero Abstract and col. 2 line 62 – col. 3 line 11 discussing the use of a model in the conversion of a NLQ to a standard query language, such as SQL. See also col. 3 lines 12-27 discussing the use of trained ML model for the conversion of a NLQ to a SQL query. It is noted that Applicants’ specification at paragraph 0034 also discusses the use of the prior art product “GPT-3”, which is a language model released in 2020.) wherein the method is performed using one or more processors. (See Romero Fig. 7 and col. 11 line 66 – col. 12 line 5 teaching an exemplary computing environment including processors and storage.)
However, Romero does not explicitly teach the remaining limitations as claimed. Haffner, though, teaches and generating the data pipeline in a programming language and applicable to a dataset based at least in part on the generated query in the standard query language, the data pipeline comprising one or more data pipeline elements, at least one data pipeline element of the one or more data pipeline elements corresponding to a query component of the generated query in the standard query language; (See Haffner page 4 , Figure 2 showing a generated QEP pipeline comprised of three component pipelines [sub-pipelines] and generated from the SQL code of “Listing 2”. See also, page 4 [section entitled “Projection”] discussing the compilation of a pipeline, and the accessing of attribute or aggregate values.)
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains to apply the teachings of Haffner for the benefit of Romero, because to do so provided a designer with options for implementing a system to dissect a query execution plan into linear sequences of operators that process tuples without the need for intermediate materialization, as taught by Haffner in the 3rd paragraph of page 4. These references were all applicable to the same field of endeavor, i.e., query processing.
Regarding claim 25: Romero teaches wherein the one or more constraints associated with the target dataset include at least one selected from a group consisting of a metric associated with the target dataset, a parameter associated with the target dataset, a parameter range associated with the target dataset, and a data range. (See Romero Fig. 1, esp. #112, showing a variety of constraints associated with the data set, including a “time” metric and other parameters.)
Regarding claim 26: Romero teaches further comprising: identifying under-specified information associated with the NL query, the under- specified information including at least one selected from a group consisting of missing information, mismatched information, a missing concept, and a mismatched concept; and generating a question based on the under-specified information. (See Romero col. 2 line 62 – col. 3 line 4 discussing disambiguating drug names / health conditions into codes or other ontologies [i.e., mismatched concepts] in order to compose the final query.)
Regarding claim 27: Romero teaches wherein the receiving a model result generated based on the NL query includes: receiving an explanation associated with the generated question; generating a model query based at least in part on the NL query and the explanation; providing the model query to the one or more computing models; and receiving the model result generated using the one or more computing models based at least in part on the model query. (See Romero col. 7 lines 43-50 and col. 8 lines 6-10 teaching that edits/updates may be supplied to the system, in the context of col. 7 lines 60-66 teaching that edits/changes may be used to update a model.)
Regarding claim 28: Romero teaches further comprising: generating a model query based on at least one selected from a group consisting of the NL query, one or more input datasets, and the target dataset; and providing the model query to the one or more computing models. (See Romero col. 2 lines 29-61 teaching the selection of pre-defined SQL subquery templates, i.e., different trained models, when converting from NLQ to SQL queries, in the context of Fig. 1 #116, 118, 120 and 122.)
Regarding claim 29: Romero teaches further comprising: selecting the one or more computing models from a set of computing models based on at least one selected from a group consisting of the NL query, one or more input datasets, and the target dataset. (See Romero col. 2 lines 29-61 teaching the selection of pre-defined SQL subquery templates, i.e., different trained models, when converting from NLQ to SQL queries, in the context of Fig. 1 #116, 118, 120 and 122.)
Regarding claim 30: Romero teaches wherein the NL query is received from a user, wherein the method further comprises: identifying an access permission associated with the target dataset; receiving permission information associated with the user; and evaluating whether the user is permitted to access the target dataset based on the access permission associated with the target dataset and the permission information associated with the target dataset. (See Romero col. 8 lines 23-29 and col. 9 line 62 – col. 10 line 7 discussing a mechanism to gain access to the database via the setting/entering of access credentials information.)
Regarding claim 31: Romero teaches further comprising: in response to the user being not permitted to access the target dataset, denying a response to the NL query. (See Romero col. 8 line 62 – col. 9 line 7 discussing denial of an access request when a user is not authorized access.)
Regarding claim 32: Romero does not explicitly teach the remaining limitations as claimed. Haffner, though, teaches further comprising: generating a query execution plan based at least in part on the generated query in the standard query language, wherein the query execution plan comprises an order of query operations, wherein the data pipeline is generated based on the query execution plan. (See Haffner page 4 paragraphs 1-4 in the context of page 4 Figure 2 and Listing 2 teaching the generation of a QEP reflecting the SQL query of Listing 2 and comprised of 3 query operations/components. Each query operation is represented a sub-pipeline, and the entire [3-part] QEP results in the complete data pipeline for the SQL/code of Listing 2.)
Regarding claim 33: Romero teaches wherein the model result includes a confidence score associated with the generated query in the standard query language. (See Romero col. 4 lines 32-49 discussing the use of a confidence score in choosing argument placeholders that are to be reflected in the SQL query. See also col. 5 lines 15-25 teaching that this modified NLQ is subsequently input into an NLQ to SQL model.)
Regarding claim 35: Romero teaches wherein the data pipeline uses one or more platform-specific expressions associated with a platform. (See Romero col. 5 lines 25-60 teaching the use of templates associated with differing ontology codes that will be reflected in the final SQL query, which in turn is reflected in the resultant data pipeline.)
Claims 36-41 are substantially similar to claims 24-29, respectively, and therefore likewise rejected.
Claim 42 is substantially similar to claim 31, and therefore likewise rejected.
Claim 43 is substantially similar to claim 24, and therefore likewise rejected.
Claim 34 is rejected under 35 U.S.C. §103 as being unpatentable over Romero Calvo et al (US Patent No. 12,124,440, hereafter referred to as “Romero”) in view of Immanuel Haffner et al. (“Fast Compilation and Execution of SQL Queries with WebAssembly”, arXiv, Cornell University Archive, document no: arXiv:2104.15098v2 [cs.DB], 3 May 2021, downloaded from: https://arxiv.org/abs/2104.15098, pp. 1-14, hereafter referred to as “Haffner”) and Clark et al (US Patent No. 7,299,237, hereafter referred to as “Clark”).
Regarding claim 34: Romero in view of Haffner does not explicitly teach the remaining limitations as claimed. Clark, though, teaches further comprising: applying the data pipeline to one or more input datasets to generate an output dataset; wherein the output dataset has a data schema that is the same as a data schema of the target dataset. (See Clark Fig. 2 in the context of the Abstract, col. 2 lines 6-12 and col. 5 lines 35-51 teaching that a data pipeline converts data into that conforming to a target schema.)
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains to apply the teachings of Clark for the benefit of Romero in view of Haffner, because to do so provided a designee with options to implement a system to dynamically and automatically transform data into a conforming format, as taught by Clark in col. 2 lines 6-12. These references were all applicable to the same field of endeavor, i.e., data conversion mechanisms.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Relevance is provided in at least the Abstract of each cited document.
Non-Patent Literature
Nguyen, Giang, et al., “Machine Learning and Deep Learning frameworks and libraries for large-scale data mining: a survey”, Artificial Intelligence Review, Volume 52, January 19, 2019, pp. 77-124.
Spark MLlib contains old RDD-based API (Resilient Distributed Dataset). RDD is the Spark’s basic abstraction of data representing an immutable, partitioned collection of elements that can be operated on in parallel with a low-level API that offers transformations and actions. Spark ML contains new API build around DataFrame-based API and ML pipelines and it is currently the primary ML API for Spark. A DataFrame is a Dataset organised into named columns and it is conceptually equivalent to a table in a relational database. Transformations and actions over a DataFrame can be specified as SQL queries, which is convenient for developers with SQL background. Moreover, Spark SQL provides to Spark more information about the structure of both the data and the computation being performed than Spark RDDAPI. SparkML brings the concept of ML pipelines, which help users to create and tune practical ML pipelines; it standardises APIs for ML algorithms so multiple ML algorithms can be combined into a single pipeline, or workflow. The MLlib is now in maintenance mode and the primary Machine Learning API for Spark is the ML package. (page 113, 2nd paragraph of section”4.3.2 Apache Spark MLlib and Spark ML”).
US Patent Application Publications
Heirich 2012/0102507
Compiler 510 may include hardware or a combination of hardware and software that may perform a source-to-source translation from a pipeline description language of input file 120 to a C-based programming language. For example, compiler 510 may read input file 120 and generate a pipeline class (e.g., OpenCL, DirectCompute, etc.) in a C++ file with a .hpp extension. (para 0054).
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to examiner ROBERT STEVENS whose telephone number is (571) 272-4102. The examiner can normally be reached Mon - Fri 6:00 - 2:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amy Ng can be reached on (571) 270-1698. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ROBERT STEVENS/Primary Examiner, Art Unit 2164
August 25, 2026