DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-5, 8-12 and 15-19 are rejected under 35 U.S.C. 103 as being unpatentable over Paulraj et al., US 12,579,124 B2 (hereinafter “Paulraj”) in view of FREILICH et al., US 2022/0019366 A1 (hereinafter “Freilich”).
Claim 1: Paulraj teaches a system comprising: at least one hardware processor; a non-transitory computer-readable medium storing instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform operations comprising:
receiving, in a first software framework, first data from a data lake, wherein the data lake stores the first data in a raw data storage format (Paulraj, [Fig. 1], [Col. 3 Lines 48-49] note data collection 110 refers to the acquisition of relevant, raw data from various data sources 101, [Col. 4 Line 27] note a data lake);
in response to a determination to store the first data in a first format object storage (Paulraj, [Col. 4 Lines 25-28] note During data preparation 120, the data is stored within a flattened data storage repository 103. Frequently, the flattened data storage repository 103 may be a data lake, data warehouse, or combination thereof):
loading the first data in an inbound buffer within a first table in the first format object storage (Paulraj, [Fig. 2], [Col. 6 Lines 38-43] note an ingestion process 250 and an output process 260 may include a configuration driven dynamic schema evolution (CDDSE) 211. The CDDSE 211 includes a dynamic schema evolution that performs a dynamic merge of different data type collections to create a merged data type collection);
performing one or more postprocessing operations on the first data and other data in the first table (Paulraj, [Fig. 4], [Col. 8 Lines 8-9] note FIG. 4 is a block diagram illustrating the components of the CDDSE 211 in greater detail, [Col. 8 Lines 38-48] note When the data produced by the data sources 101 is saved in the Delta format, data acquired from a data source 101 at different times or from different combinations of data sources 101 may be merged together using an automatic schema evolution to create a data table. For example, as the data from the data sources 101 is saved in the Delta format, the data may be saved in columns, if a column in the data from the data sources 101 is not available in the previously saved flattened table, a new column may be added to the table that is updated with the values in the data from the data sources 101); and
merging the postprocessed first data and the other data into an active data portion of the first table (Paulraj, [Col. 5 Lines 30-35] note in the deployment stage 140, the output from the models 105 are provided to application integrations 107. As illustrated, the output from models 105 can be provided to multiple application integrations 107. The application integrations are referred to herein generally or collectively as application integrations 107).
Paulraj does not explicitly teach providing read access but not write access to the postprocessed first data and the other data in the first table to one or more processes of the first software framework.
However, Freilich teaches this (Frielich, [0198] note storage systems described herein may be used to form a data lake. A data lake may operate as the first place that an organization's data flows to, where such data may be in a raw format, [0267] note data services are not provided for the dataset 730 until each of the multiple hosts has redirected its access to the dataset 730 from the source storage system 716 to the target storage system 706. In some implementations, multipathing may be utilized, during a transition period, to allow the host 740 to access the dataset 730 through either the source storage system 716 (for read access only) or the target storage system 706, [0274] note any read-write access request or data services request targeting a logical address of the migrated data blocks 832, 834 will be serviced on the storage resources 810 of the target storage system 806).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the deployment environment of Paulraj with the data access services of Freilich according to known methods (i.e. providing access to a dataset during a transition). Motivation for doing so is that it would be advantageous to provide a storage system that can manage the migration without making the data unavailable during the migration and that can perform the migration without using host resources and networks (Freilich, [0257]).
Claim 2: Paulraj and Freilich teach the system of claim 1, wherein the operations further comprise:
transferring the postprocessed first data to a second table; and providing both read access and write access to the postprocessed first data in the second table to the one or more processes of the first software framework (Freilich, [0274] note metadata objects 864, 866 have been updated to point to data blocks 832, 834 in the storage resources 810 of the target storage system, while metadata objects 868, 870 still point to unmigrated data blocks 836, 838 in the source storage system 816. Thus, any read-write access request or data services request targeting a logical address of the migrated data blocks 832, 834 will be serviced on the storage resources 810 of the target storage system 806).
Claim 3: Paulraj and Freilich teach the system of claim 1, wherein the first format object storage is an open data format object storage, and wherein the first table is a delta table (Paulraj, [Col. 10 Lines 32-37] note method 500 proceeds at 507, where a flattened enriched Delta table is generated. Further, the method 500 proceeds at 509 where a merge schema option is enabled for the enriched Delta table. An exemplary flattened enriched Delta table 533 is illustrated in FIG. 5B).
Claim 4: Paulraj and Freilich teach the system of claim 1, wherein the inbound buffer is implemented as a Parquet file (Paulraj, [Col. 8 Lines 21-34] note he output from the models may be saved in particular formats like Avro, Parquet, and Delta, which are file formats used for storing and processing large-scale data in big data systems. Avro is a row-based file format that uses a compact, binary data serialization format. It is designed to be language-independent, thus data written in one programming language can be read in another language. Parquet is a columnar storage file format that is designed to be highly efficient for processing large datasets, as it allows for parallel processing of individual columns and can compress data for better storage utilization. Delta is a transaction storage layer built on top of Parquet that is designed for building data lakes and data warehouses).
Claim 5: Paulraj and Freilich teach the system of claim 1, wherein the performing the one or more postprocessing operations is performed periodically based on a set period (Paulraj, [Col. 18 Lines 23-25] note In a moderately changing model 803, the models may be created periodically or in response to events that occur at a moderately rate).
Claim 8: Paulraj teaches a method comprising:
receiving, in a first software framework, first data from a data lake, wherein the data lake stores the first data in a raw data storage format (Paulraj, [Fig. 1], [Col. 3 Lines 48-49] note data collection 110 refers to the acquisition of relevant, raw data from various data sources 101, [Col. 4 Line 27] note a data lake);
in response to a determination to store the first data in a first format object storage (Paulraj, [Col. 4 Lines 25-28] note During data preparation 120, the data is stored within a flattened data storage repository 103. Frequently, the flattened data storage repository 103 may be a data lake, data warehouse, or combination thereof):
loading the first data in an inbound buffer within a first table in the first format object storage (Paulraj, [Fig. 2], [Col. 6 Lines 38-43] note an ingestion process 250 and an output process 260 may include a configuration driven dynamic schema evolution (CDDSE) 211. The CDDSE 211 includes a dynamic schema evolution that performs a dynamic merge of different data type collections to create a merged data type collection);
performing one or more postprocessing operations on the first data and other data in the first table (Paulraj, [Fig. 4], [Col. 8 Lines 8-9] note FIG. 4 is a block diagram illustrating the components of the CDDSE 211 in greater detail, [Col. 8 Lines 38-48] note When the data produced by the data sources 101 is saved in the Delta format, data acquired from a data source 101 at different times or from different combinations of data sources 101 may be merged together using an automatic schema evolution to create a data table. For example, as the data from the data sources 101 is saved in the Delta format, the data may be saved in columns, if a column in the data from the data sources 101 is not available in the previously saved flattened table, a new column may be added to the table that is updated with the values in the data from the data sources 101); and
merging the postprocessed first data and the other data into an active data portion of the first table (Paulraj, [Col. 5 Lines 30-35] note in the deployment stage 140, the output from the models 105 are provided to application integrations 107. As illustrated, the output from models 105 can be provided to multiple application integrations 107. The application integrations are referred to herein generally or collectively as application integrations 107); and
Paulraj does not explicitly teach providing read access but not write access to the postprocessed first data and the other data in the first table to one or more processes of the first software framework.
However, Freilich teaches this (Frielich, [0198] note storage systems described herein may be used to form a data lake. A data lake may operate as the first place that an organization's data flows to, where such data may be in a raw format, [0267] note data services are not provided for the dataset 730 until each of the multiple hosts has redirected its access to the dataset 730 from the source storage system 716 to the target storage system 706. In some implementations, multipathing may be utilized, during a transition period, to allow the host 740 to access the dataset 730 through either the source storage system 716 (for read access only) or the target storage system 706, [0274] note any read-write access request or data services request targeting a logical address of the migrated data blocks 832, 834 will be serviced on the storage resources 810 of the target storage system 806).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the deployment environment of Paulraj with the data access services of Freilich according to known methods (i.e. providing access to a dataset during a transition). Motivation for doing so is that it would be advantageous to provide a storage system that can manage the migration without making the data unavailable during the migration and that can perform the migration without using host resources and networks (Freilich, [0257]).
Claim 9: Paulraj and Freilich teach the method of claim 8, further comprising:
transferring the postprocessed first data to a second table; and providing both read access and write access to the postprocessed first data in the second table to the one or more processes of the first software framework (Freilich, [0274] note metadata objects 864, 866 have been updated to point to data blocks 832, 834 in the storage resources 810 of the target storage system, while metadata objects 868, 870 still point to unmigrated data blocks 836, 838 in the source storage system 816. Thus, any read-write access request or data services request targeting a logical address of the migrated data blocks 832, 834 will be serviced on the storage resources 810 of the target storage system 806).
Claim 10: Paulraj and Freilich teach the method of claim 8, wherein the first format object storage is an open data format object storage, and wherein the first table is a delta table (Paulraj, [Col. 10 Lines 32-37] note method 500 proceeds at 507, where a flattened enriched Delta table is generated. Further, the method 500 proceeds at 509 where a merge schema option is enabled for the enriched Delta table. An exemplary flattened enriched Delta table 533 is illustrated in FIG. 5B).
Claim 11: Paulraj and Freilich teach the method of claim 8, wherein the inbound buffer is implemented as a Parquet file (Paulraj, [Col. 8 Lines 21-34] note he output from the models may be saved in particular formats like Avro, Parquet, and Delta, which are file formats used for storing and processing large-scale data in big data systems. Avro is a row-based file format that uses a compact, binary data serialization format. It is designed to be language-independent, thus data written in one programming language can be read in another language. Parquet is a columnar storage file format that is designed to be highly efficient for processing large datasets, as it allows for parallel processing of individual columns and can compress data for better storage utilization. Delta is a transaction storage layer built on top of Parquet that is designed for building data lakes and data warehouses).
Claim 12: Paulraj and Freilich teach the method of claim 8, wherein the performing the one or more postprocessing operations is performed periodically based on a set period (Paulraj, [Col. 18 Lines 23-25] note In a moderately changing model 803, the models may be created periodically or in response to events that occur at a moderately rate).
Claim 15: Paulraj teaches a non-transitory machine-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform operations comprising:
receiving, in a first software framework, first data from a data lake, wherein the data lake stores the first data in a raw data storage format (Paulraj, [Fig. 1], [Col. 3 Lines 48-49] note data collection 110 refers to the acquisition of relevant, raw data from various data sources 101, [Col. 4 Line 27] note a data lake);
in response to a determination to store the first data in a first format object storage (Paulraj, [Col. 4 Lines 25-28] note During data preparation 120, the data is stored within a flattened data storage repository 103. Frequently, the flattened data storage repository 103 may be a data lake, data warehouse, or combination thereof):
loading the first data in an inbound buffer within a first table in the first format object storage (Paulraj, [Fig. 2], [Col. 6 Lines 38-43] note an ingestion process 250 and an output process 260 may include a configuration driven dynamic schema evolution (CDDSE) 211. The CDDSE 211 includes a dynamic schema evolution that performs a dynamic merge of different data type collections to create a merged data type collection);
performing one or more postprocessing operations on the first data and other data in the first table (Paulraj, [Fig. 4], [Col. 8 Lines 8-9] note FIG. 4 is a block diagram illustrating the components of the CDDSE 211 in greater detail, [Col. 8 Lines 38-48] note When the data produced by the data sources 101 is saved in the Delta format, data acquired from a data source 101 at different times or from different combinations of data sources 101 may be merged together using an automatic schema evolution to create a data table. For example, as the data from the data sources 101 is saved in the Delta format, the data may be saved in columns, if a column in the data from the data sources 101 is not available in the previously saved flattened table, a new column may be added to the table that is updated with the values in the data from the data sources 101); and
merging the postprocessed first data and the other data into an active data portion of the first table (Paulraj, [Col. 5 Lines 30-35] note in the deployment stage 140, the output from the models 105 are provided to application integrations 107. As illustrated, the output from models 105 can be provided to multiple application integrations 107. The application integrations are referred to herein generally or collectively as application integrations 107); and
Paulraj does not explicitly teach providing read access but not write access to the postprocessed first data and the other data in the first table to one or more processes of the first software framework.
However, Freilich teaches this (Frielich, [0198] note storage systems described herein may be used to form a data lake. A data lake may operate as the first place that an organization's data flows to, where such data may be in a raw format, [0267] note data services are not provided for the dataset 730 until each of the multiple hosts has redirected its access to the dataset 730 from the source storage system 716 to the target storage system 706. In some implementations, multipathing may be utilized, during a transition period, to allow the host 740 to access the dataset 730 through either the source storage system 716 (for read access only) or the target storage system 706, [0274] note any read-write access request or data services request targeting a logical address of the migrated data blocks 832, 834 will be serviced on the storage resources 810 of the target storage system 806).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the deployment environment of Paulraj with the data access services of Freilich according to known methods (i.e. providing access to a dataset during a transition). Motivation for doing so is that it would be advantageous to provide a storage system that can manage the migration without making the data unavailable during the migration and that can perform the migration without using host resources and networks (Freilich, [0257]).
Claim 16: Paulraj and Freilich teach the non-transitory machine-readable medium of claim 15, wherein the operations further comprise:
transferring the postprocessed first data to a second table; and providing both read access and write access to the postprocessed first data in the second table to the one or more processes of the first software framework (Freilich, [0274] note metadata objects 864, 866 have been updated to point to data blocks 832, 834 in the storage resources 810 of the target storage system, while metadata objects 868, 870 still point to unmigrated data blocks 836, 838 in the source storage system 816. Thus, any read-write access request or data services request targeting a logical address of the migrated data blocks 832, 834 will be serviced on the storage resources 810 of the target storage system 806).
Claim 17: Paulraj and Freilich teach the non-transitory machine-readable medium of claim 15, wherein the first format object storage is an open data format object storage, and wherein the first table is a delta table (Paulraj, [Col. 10 Lines 32-37] note method 500 proceeds at 507, where a flattened enriched Delta table is generated. Further, the method 500 proceeds at 509 where a merge schema option is enabled for the enriched Delta table. An exemplary flattened enriched Delta table 533 is illustrated in FIG. 5B).
Claim 18: Paulraj and Freilich teach the non-transitory machine-readable medium of claim 15, wherein the inbound buffer is implemented as a Parquet file (Paulraj, [Col. 8 Lines 21-34] note he output from the models may be saved in particular formats like Avro, Parquet, and Delta, which are file formats used for storing and processing large-scale data in big data systems. Avro is a row-based file format that uses a compact, binary data serialization format. It is designed to be language-independent, thus data written in one programming language can be read in another language. Parquet is a columnar storage file format that is designed to be highly efficient for processing large datasets, as it allows for parallel processing of individual columns and can compress data for better storage utilization. Delta is a transaction storage layer built on top of Parquet that is designed for building data lakes and data warehouses).
Claim 19: Paulraj and Freilich teach the non-transitory machine-readable medium of claim 15, wherein the performing the one or more postprocessing operations is performed periodically based on a set period (Paulraj, [Col. 18 Lines 23-25] note In a moderately changing model 803, the models may be created periodically or in response to events that occur at a moderately rate).
Claims 6, 7, 13, 14 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Paulraj and Freilich in further view of Oattes et al., US 2023/0359614 A1 (hereinafter “Oattes”).
Claim 6: Paulraj and Freilich do not explicitly teach the system of claim 1, wherein the performing the one or more postprocessing operations is performed when a maximum latency of the inbound buffer is reached.
However, Oattes teaches this (Oattes, [0007] note data lakes, [0268] note RDF data can be ingested by the DPA into a CADS staging table temporarily, [0255] note a maximum batch size may be preconfigured by the DPA. When an in-memory list of triples reaches this limit within the loading service, the data is inserted into the schema tables, [0270] note the DPA can buffer incoming queries, up to a specific batch size (number of queries) and/or a maximum latency (5 minutes) and only process that batch when a “lull” in the input is detected).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the deployment environment of Paulraj and Freilich with the batch data ingestion of Oattes according to known methods (i.e. ingesting batches of data based on a maximum batch size and/or maximum latency). Motivation for doing so is that this provides an optimized batch load capability (Oattes, [0255]).
Claim 7: Paulraj and Freilich do not explicitly teach the system of claim 1, wherein the performing the one or more postprocessing operations is performed when a maximum size of the inbound buffer is reached.
However, Oattes teaches this (Oattes, [0007] note data lakes, [0268] note RDF data can be ingested by the DPA into a CADS staging table temporarily, [0255] note a maximum batch size may be preconfigured by the DPA. When an in-memory list of triples reaches this limit within the loading service, the data is inserted into the schema tables, [0270] note the DPA can buffer incoming queries, up to a specific batch size (number of queries) and/or a maximum latency (5 minutes) and only process that batch when a “lull” in the input is detected).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the deployment environment of Paulraj and Freilich with the batch data ingestion of Oattes according to known methods (i.e. ingesting batches of data based on a maximum batch size and/or maximum latency). Motivation for doing so is that this provides an optimized batch load capability (Oattes, [0255]).
Claim 13: Paulraj and Freilich do not explicitly teach the method of claim 8, wherein the performing the one or more postprocessing operations is performed when a maximum latency of the inbound buffer is reached.
However, Oattes teaches this (Oattes, [0007] note data lakes, [0268] note RDF data can be ingested by the DPA into a CADS staging table temporarily, [0255] note a maximum batch size may be preconfigured by the DPA. When an in-memory list of triples reaches this limit within the loading service, the data is inserted into the schema tables, [0270] note the DPA can buffer incoming queries, up to a specific batch size (number of queries) and/or a maximum latency (5 minutes) and only process that batch when a “lull” in the input is detected).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the deployment environment of Paulraj and Freilich with the batch data ingestion of Oattes according to known methods (i.e. ingesting batches of data based on a maximum batch size and/or maximum latency). Motivation for doing so is that this provides an optimized batch load capability (Oattes, [0255]).
Claim 14: Paulraj and Freilich do not explicitly teach the method of claim 8, wherein the performing the one or more postprocessing operations is performed when a maximum size of the inbound buffer is reached.
However, Oattes teaches this (Oattes, [0007] note data lakes, [0268] note RDF data can be ingested by the DPA into a CADS staging table temporarily, [0255] note a maximum batch size may be preconfigured by the DPA. When an in-memory list of triples reaches this limit within the loading service, the data is inserted into the schema tables, [0270] note the DPA can buffer incoming queries, up to a specific batch size (number of queries) and/or a maximum latency (5 minutes) and only process that batch when a “lull” in the input is detected).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the deployment environment of Paulraj and Freilich with the batch data ingestion of Oattes according to known methods (i.e. ingesting batches of data based on a maximum batch size and/or maximum latency). Motivation for doing so is that this provides an optimized batch load capability (Oattes, [0255]).
Claim 20: Paulraj and Freilich do not explicitly teach the non-transitory machine-readable medium of claim 15, wherein the performing the one or more postprocessing operations is performed when a maximum latency of the inbound buffer is reached.
However, Oattes teaches this (Oattes, [0007] note data lakes, [0268] note RDF data can be ingested by the DPA into a CADS staging table temporarily, [0255] note a maximum batch size may be preconfigured by the DPA. When an in-memory list of triples reaches this limit within the loading service, the data is inserted into the schema tables, [0270] note the DPA can buffer incoming queries, up to a specific batch size (number of queries) and/or a maximum latency (5 minutes) and only process that batch when a “lull” in the input is detected).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the deployment environment of Paulraj and Freilich with the batch data ingestion of Oattes according to known methods (i.e. ingesting batches of data based on a maximum batch size and/or maximum latency). Motivation for doing so is that this provides an optimized batch load capability (Oattes, [0255]).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Giuseppi Giuliani whose telephone number is (571)270-7128. The examiner can normally be reached Monday-Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kavita Stanley can be reached at (571)272-8352. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/GIUSEPPI GIULIANI/Primary Examiner, Art Unit 2153