DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
The amendment filed 6/8/2026 has been entered. Claims 11-20 stand amended. Claims 1-20 stand pending.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Oswal et al. in US Patent Application Publication № 2024/0028605, hereinafter called Oswal, in combination with Kalmanek, Jr. et. al. in US Patent Application Publication № 2023/0394043, hereinafter called Kalmanek.
In regard to claim 1, Oswal teaches a method comprising:
storing, by a first computing node of a set of computing nodes of a plurality of computing nodes of a parallelized database system (“The DBMS executing on a distributed cluster of nodes may add additional factors, such as the number of nodes and/or the size of memory on each node ( especially the coordinator node).” Paragraph 0021), a plurality of data format options in local memory (“In an embodiment, the encoding of the data format may be different based on where the data is stored: volatile memory or persistent storage. However, once loaded into memory for operation, the data portion's encoding may be changed or preserved depending on the benefits of the operations on the data portion” paragraph 0018);
formatting, by a lead computing node of the set of computing nodes, the dataset in a primary data format to produce a primary data formatted dataset (i.e. OF, “The format that corresponds to the data's original format (e.g., the on-disk/persistent storage format) is referred to herein as the "original format" or "OF". Data that is in the original format is referred to herein as OF data. For example, string data of a column may be originally stored encoded in the variable-length format (VARLEN). The alternative format in which the data portion may be loaded into memory (e.g., volatile memory) is referred to herein as a "mirror format" or "MF". Data that is encoded in the mirror format is referred to herein as MF data” paragraph 0019);
storing, by the lead computing node, the primary data formatted dataset in system state data, wherein the system state data is stored in memory of at least one computing node of the set of computing nodes (“The format that corresponds to the data's original format (e.g., the on-disk/persistent storage format) is referred to herein as the "original format" or "OF".” Paragraph 0019), and wherein the system state data is mediated by the set of computing nodes in accordance with a consensus protocol (i.e. offload partition selections, paragraph 0043);
receiving, by the first computing node, a query request involving the dataset, wherein the query request includes a first data function (“The workload set may represent the most typical set of queries that are executed by the DBMS. For example, in the distributed DBMS, such a workload set may include OLAP (Online Analytical Processing) queries. The OLAP queries are generally read oriented queries that may include a large number of scans, base relation (e.g., join operation) and sort operations with predicate evaluation(s) on large data sets” paragraph 0016);
determining, by the first computing node, that the first data function triggers a first data conversion optimization (“The workload set may reference various data portions (e.g., columns), which are referred to herein as "candidate data portions." A candidate data portion may be stored in DBMS in a particular encoding format having its own execution benefits. However, these execution benefits may not match the usage of the candidate data portion in the workload set.” Paragraph 0017);
accessing, by the first computing node, the primary data formatted dataset from the system state data (“An operation on MF data, as opposed to OF data, may execute longer or shorter due to de-compression and network delays required by operation(s) of the query in the workload.” Paragraph 0020, wherein “Accordingly, various factors, such as execution statistics and data statistics of a data portion in the query operations, are used to determine whether to encode the data portion in the MF data or use OF data” paragraph 0021);
converting, by the first computing node, the primary data formatted dataset from the primary data format to a secondary data format of the plurality of data format options in accordance with the first data conversion optimization to produce a secondary data formatted dataset (“However, once loaded into memory for operation, the data portion's encoding may be changed or preserved depending on the benefits of the operations on the data portion.” Paragraph 0018);
and processing, by the first computing node, the secondary data formatted dataset in accordance with the first data function to produce a first function output (“After selecting an optimal execution plan for a query, coordinator node 150 obtains and executes the execution plan, in an embodiment” paragraph 0039).
However, Oswal fails to expressly teach ingesting, by the set of computing nodes, a dataset for storage within the parallelized database system.
Kalmanek teaches ingesting, by the set of computing nodes, a dataset for storage within the parallelized database system (“The system ingests data from various data lake sources and partition the data into multiple data partitions” paragraph 0025).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the instant invention to modify the distributed database system which stores data in multiple formats, as taught by Oswal (in at least paragraph 0021), to include the ingestion of data to be divided into multiple formats, as taught by Kalamanek (in at least paragraphs 0025 and 0049). One would have been motivated to do so in order to allow a data lake to grow, as taught by Kalamanek in at least paragraph 0024, and to allow partitioning to change in response to change sin the underlying data, as taught by Kalamanek in at least paragraph 0076.
In regard to claim 11, it is substantially similar to claim 1, and accordingly is rejected under similar reasoning.
In regard to claim 2, Oswal further teaches that the formatting the dataset in the primary data format comprises:
selecting, by the lead computing node, the primary data format based on one or more of:
applying, by the lead computing node, a default setting indicating the primary data format (“For example, one data portion may be persistently maintained in one format (and in some embodiments, only in that format),” paragraph 0018);
obtaining, by the lead computing node, a user input indicating the primary data format (paragraph 0138);
accessing, by the lead computing node, configuration data for the system state data, wherein the configuration data indicates the primary data format (i.e. machine learning model, “In an embodiment, the DBMS may include multiple machine learning models, each for different query operators (e.g., scan, decode, partitioning, group-by), for different costs (e.g., network delay, memory size), and for encoding type and/or combinations thereof. For example, a machine learning model may be trained to determine a base relation operation's network cost when altering from OF data to MF data.” Paragraph 0029);
performing, by the lead computing node, an algorithm use analysis of at least a portion of the parallelized database system to determine that the primary data format is used most often in algorithms performed by the at least a portion of the parallelized database system (“The DBMS analyzes each of the queries to obtain the data statistics for the data portion(s) referenced in the operation(s) of the queries. The DBMS accesses the data and pre-execution statistics regarding the relevant data portions and generates training data set of features” paragraph 0023);
and performing, by the lead computing node, a cost analysis to determine the primary data format (“Techniques are described for analyzing a query workload set to estimate the benefit/cost measurement of the workload set execution when candidate data portion(s) encoding format is altered to an alternative encoded format when used in the execution.” Paragraph 0018).
In regard to claim 12, it is substantially similar to claim 2, and accordingly is rejected under similar reasoning.
In regard to claim 3, Oswal further teaches that the determining that the first data function triggers the first data conversion optimization comprises:
applying, by the first computing node, a default setting indicating that the secondary data format is optimal for processing the first data function (“For example, one data portion may be persistently maintained in one format (and in some embodiments, only in that format),” paragraph 0018);
obtaining, by the first computing node, a user input indicating the secondary data format (paragraph 0138);
accessing, by the first computing node, local configuration data, wherein the local configuration data indicates that the secondary data format is optimal for processing the first data function (i.e. machine learning model, “In an embodiment, the DBMS may include multiple machine learning models, each for different query operators (e.g., scan, decode, partitioning, group-by), for different costs (e.g., network delay, memory size), and for encoding type and/or combinations thereof. For example, a machine learning model may be trained to determine a base relation operation's network cost when altering from OF data to MF data.” Paragraph 0029);
performing, by the first computing node, an analysis of the first data function to determine that the secondary data format is optimal for use in processing the first function used (“A trained machine learning model for an operator may receive as input the feature set values of data portion(s) (same as the feature set used in the training) and generates an estimate of performance gain for altering the encoding of the data portions.” Paragraph 0025, note that the feature values are stored, paragraph 0028);
and performing, by the first computing node, a cost analysis to determine that the secondary data format is more cost effective for processing the first data function than the primary data format (“A trained machine learning model for an operator may receive as input the feature set values of data portion(s) (same as the feature set used in the training) and generates an estimate of performance gain for altering the encoding of the data portions.” Paragraph 0025).
In regard to claim 13, it is substantially similar to claim 3, and accordingly is rejected under similar reasoning.
In regard to claim 4, Oswal further teaches that the accessing the primary data formatted dataset from the system state data comprises one of: obtaining, by the first computing node, a locally stored copy of the primary data formatted dataset;
and retrieving, by the first computing node, the primary data formatted dataset from one or more other computing nodes of the set of computing nodes (i.e. by selecting to offload or not, paragraph 0041, note that the format data is included in this decision, as in at least paragraph 0040).
In regard to claim 14, it is substantially similar to claim 4, and accordingly is rejected under similar reasoning.
In regard to claim 5, Oswal further teaches: storing, by the first computing node, the secondary data formatted dataset in local memory (“The alternative format in which the data portion may be loaded into memory (e.g., volatile memory) is referred to herein as a "mirror format" or "MF". Data that is encoded in the mirror format is referred to herein as MF data. For example, string data of a column may be mirrored in memory in the dictionary encoding format.” Paragraph 0019).
In regard to claim 15, it is substantially similar to claim 5, and accordingly is rejected under similar reasoning.
In regard to claim 6, Kalmanek further teaches
obtaining, by the first computing node, a data update to the primary data formatted dataset in the system state data, wherein the data update was mediated by the set of computing nodes (“For example, the ingestion engine 108 may obtain data from a data lake source every hour, every day, every week, every month, or other suitable time period. In some embodiments, the frequency of obtaining data from a data lake source may be a configurable parameter that may be adjusted by a user (e.g., administrator) of the data lake 100.” Paragraph 0042), and wherein the data update was applied to the primary data formatted dataset to produce an updated primary data formatted dataset (i.e. a first format updated); and updating, by the first computing node, the secondary data formatted dataset based on the updated primary data formatted dataset to produce an updated secondary data formatted dataset (“In some embodiments, the data partitions 106 may be configured to store data in multiple storage formats. For example, the data partitions 106 may store a portion of the data in a columnar format and another portion of the data in a row oriented format. The data partitions 106 may store different portions of data in different formats to optimize query performance on the respective portions of data. For example, the data partitions 106 may: (1) store a first portion of data in a columnar format because queries aggregating data from multiple different records are more likely to be executed on the first portion of data; and (2) store a second portion of data in a row oriented format because queries for specific records are more likely to be executed on the second portion of data” paragraph 0056, wherein “In some embodiments, the ingestion engine 108 may be configured to determine one or more values to index the partitions with based on how the data is expected to be queried. For example, the ingestion engine 108 may index the data on a particular field based on determining that the field is likely to be queried by a user of the data lake 100.” Paragraph 0051).
In regard to claim 16, it is substantially similar to claim 6, and accordingly is rejected under similar reasoning.
In regard to claim 7, Oswal further teaches that:
the query request further includes a second data function; determining, by the first computing node, that the second data function does not trigger a second data conversion optimization (i.e. is better in the OF, “The total estimated
performance gain and cost may determine whether the candidate column is to be encoded in the dictionary encoding or left in the VARLEN string format.” Paragraph 0031, wherein the functions are taught in at least paragraph 0029);
and processing, by the first computing node, the primary data formatted dataset in accordance with the second data function to produce a second function output in parallel with processing the secondary data formatted dataset in accordance with the first data function to produce the first function output (i.e. executed by selecting data, as in paragraph 0040).
In regard to claim 17, it is substantially similar to claim 7, and accordingly is rejected under similar reasoning.
In regard to claim 8, Oswal further teaches: wherein the query request further includes a second data function (i.e. an additional operation, “The DBMS analyzes each of the queries to obtain the data statistics for the data portion(s) referenced in the operation(s) of the queries. The DBMS accesses the data and pre-execution statistics regarding the relevant data portions and generates training data set of features.” Paragraph 0023);
determining, by the first computing node, that the second data function triggers a second data conversion optimization (paragraph 0029);
converting, by the first computing node, the primary data formatted dataset from the primary data format to a third data format of the plurality of data format options in accordance with the second data conversion optimization to produce a third data formatted dataset (“To generate the result set for the training data set, the DBMS may modify the format in which one or more of the data portions referenced by the training set of queries are executed by encoding the data portions in the alternative format and storing the data portions in the memory. After the encoding, the DBMS executes the training set of queries again to collect execution time statistics. Based on the execution time statistics, the DBMS determines performance cost/gain information for each data portion and the operation(s) thereof, thereby generating result set values for each data portion and the operation(s) thereof, in an embodiment.” Paragraph 0024);
And processing, by the first computing node, the third data formatted dataset in accordance with the second data function to produce a second function output in parallel with processing the secondary data formatted dataset in accordance with the first data function to produce the first function output (i.e. execute the query plan, paragraph 0039, which may be offloaded, paragraph 0040).
In regard to claim 18, it is substantially similar to claim 8, and accordingly is rejected under similar reasoning.
In regard to claim 9, Oswal further teaches: storing, by a second computing node of the set of computing nodes, a second plurality of data format options in local memory (paragraph 0045, wherein specific examples are taught in paragraph 0046; alternatively or additionally, subsets, paragraph 0047).
In regard to claim 19, it is substantially similar to claim 9, and accordingly is rejected under similar reasoning.
In regard to claim 10, Kalmanek further teaches ingesting, by a second set of computing nodes of the plurality of computing nodes, the dataset for storage within the parallelized database system (“For example, the ingestion engine 108 may obtain data from a data lake source every hour, every day, every week, every month, or other suitable time period. In some embodiments, the frequency of obtaining data from a data lake source may be a configurable parameter that may be adjusted by a user (e.g., administrator) of the data lake 100.” Paragraph 0042);
formatting, by a lead computing node of the second set of computing nodes, the dataset in a second primary data format to produce a second primary data formatted dataset (i.e. a first format); and storing, by the lead computing node, the second primary data formatted dataset in second system state data, wherein the second system state data is stored in memory of at least one computing node of the second set of computing nodes (“In some embodiments, the data partitions 106 may be configured to store data in multiple storage formats. For example, the data partitions 106 may store a portion of the data in a columnar format and another portion of the data in a row oriented format. The data partitions 106 may store different portions of data in different formats to optimize query performance on the respective portions of data. For example, the data partitions 106 may: (1) store a first portion of data in a columnar format because queries aggregating data from multiple different records are more likely to be executed on the first portion of data; and (2) store a second portion of data in a row oriented format because queries for specific records are more likely to be executed on the second portion of data” paragraph 0056, wherein “In some embodiments, the ingestion engine 108 may be configured to determine one or more values to index the partitions with based on how the data is expected to be queried. For example, the ingestion engine 108 may index the data on a particular field based on determining that the field is likely to be queried by a user of the data lake 100.” Paragraph 0051),
and wherein the second system state data is mediated by the second set of computing nodes in accordance with a second consensus protocol (“In such embodiments, the ingestion engine 108 may have concurrency control to safely operate on data partition(s) within the shard without locks.” Paragraph 0047).
In regard to claim 20, it is substantially similar to claim 10, and accordingly is rejected under similar reasoning.
Response to Arguments
Applicant’s arguments, see page 14, filed 6/8/2026, with respect to the rejection of claims 1-20 under 35 U.S.C. 101 have been fully considered and, in light of the instant amendments, are persuasive. The rejection of claims 11-20 under 35 U.S.C. 101 has been withdrawn.
Applicant’s arguments, see page 14-15, filed 6/8/2026, with respect to the rejection of claims 1-20 under 35 U.S.C. 103 have been fully considered and are not persuasive.
In particular, applicants’ argument that Oswal fails to teach “storing, by the lead computing node, the primary data formatted dataset in system state data, wherein the system state data is stored in memory of at least one computing node of the set of computing nodes, and wherein the system state data is mediated by the set of computing nodes in accordance with a consensus protocol” is not persuasive, as the broadest reasonable interpretation of this claim limitation is taught. Applicant appears to highlight the original format data and the encoded format data taught by Oswal as being strictly alternative to the system state data recited in the instant claim (see applicant’s remarks, page 15, first full paragraph). However, the instant claim language does not appear to require any such non-overlap in data storage (and indeed, instant claim 1 appears to expressly require that the claimed primary data formatted data set be stored as a component of the system state data, “…storing, by the lead computing node, the primary data formatted dataset in system state data,…”). Oswal does, in fact, teach a system state data, in that Oswal teaches data which defines a portion of the state of a system, and Oswal teaches that the data meets the limitations ascribed to the ‘system state data’.
First, Oswal teaches that the system state data ‘[stores], by the lead computing node, the primary data formatted dataset in system state data”, in that Oswal teaches “The format that corresponds to the data's original format (e.g., the on-disk/persistent storage format) is referred to herein as the "original format" or "OF". (paragraph 0019). Absent any further technical limitation as to the nature, structure, or use of the system state data, the instant teaching meets the broadest reasonable interpretation of the claim, as the data is related to the system state.
[Examiner’s Note: the Offload Engine taught to be part of the Coordinator node in at least paragraph 0040 is taught to manage the selections of data that are encoded into different formats, as taught in at least paragraph 0044, which, though not necessary to meet the broadest reasonable limitation of the instant claim limitations, would appear to include data of at least some form related to the current state of the system, i.e. system state data.]
Second, Oswal teaches that that “the system state data is stored in memory of at least one computing node of the set of computing nodes” (This is taught in paragraph 0019, and again more explicitly, “Coordinator node 150 includes offload engine 151 for offloading operations to be executed on worker nodes 110A/B. Work nodes 110A/B include mirror format (MF) data 116A/B and original format (OF) data 117 A/B that store one or more data portions in memory 115A/B that are also stored in OF data 157 of coordinator node 150”, paragraph 0040. The coordinator node, which is at least one computing node, is taught to store the OF data, i.e. the system state data.)
Third, Oswal teaches that “the system state data is mediated by the set of computing nodes in accordance with a consensus protocol” (“When offload engine 151 determines that an operation of an execution plan is to be offloaded onto worker nodes, offload engine 151 may invoke execution engines 114A/B of worker nodes 110A/B to execute the operation.” Paragraph 0043.; note that this includes a portion of the OF data and determination of which portions to encode, as in at least paragraph 0044) The instant teaching meets the broadest reasonable interpretation of the limitation, as a consensus, i.e. consistent global state, is achieved through a protocol, i.e. a particular way of performing a task. Accordingly, absent any further technical limitations as to the nature, algorithm, or structure of the claimed ‘consensus protocol’, the broadest reasonable interpretation of a consensus protocol is taught.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Lauren Z Ganger whose telephone number is (571)272-0270. The examiner can normally be reached 10:00 AM - 7:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ajay Bhatia can be reached at (571) 272-3906. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AJAY M BHATIA/ Supervisory Patent Examiner, Art Unit 2156