IP Library › Granted Patent US 12,530,368
Granted Patent B1
US 12,530,368 · App. 18/779,282 · Granted Jan 20, 2026

Scalable data import into managed lakehouses

Inventors: Zhou Fang (Santa Clara, CA); Jian Guo (Kirkland, WA); Thibaud Hottelier (Seattle, WA); Anoop Kochummen Johnson (Fremont, CA); Micah Kornfield (Seattle, WA); Justin Levandoski (Seattle, WA); Yuri Volobuev (Walnut Creek, CA); Yiwei Zhang (Bellevue, WA)
Assignee: Google LLC
G06F16/258G06F16/2358
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,368
App. No.
18/779,282
Granted
Jan 20, 2026
Kind
B1
Abstract

Aspects of the disclosure are directed to managing files in data lakehouses using rewrite-free loading. Rewrite-free loading includes keeping track of information from data files that could be missing when imported to the data lakehouses without having to perform full-copy loading. Rewrite-free loading can store this information in table metadata or augment headers and/or footers of the data files with this information when importing to a data lakehouse. Rewrite-free loading allows for more accurate management of data lakehouses with lower computational costs.

Claims (45)

1 . A method for managing files in a data lakehouse comprising:

receiving, by one or more processors, a request from a query engine to load a data file to the data lakehouse;

retrieving, by the one or more processors, the data file and information associated with the data file;

generating, by the one or more processors, a header/footer file containing the information associated with the data file;

generating, by the one or more processors, a composite file based on the data file and the header/footer file by concatenating the data file with the header/footer file, wherein at least one of a header or footer of the data file becomes available space in memory; and

storing, by the one or more processors, the composite file in the data lakehouse.

2 . The method of claim 1 , wherein the information associated with the data file comprises at least one of a stable column identifier, a partition key, or an integrity signature.

3 . The method of claim 1 , wherein the header or footer of the data file becomes available space in memory in response to being concatenated with the header/footer file.

4 . The method of claim 1 , further comprising storing, by the one or more processors, an extension in the available space in memory between the data file and the header/footer file of the composite file.

5 . The method of claim 4 , wherein the extension comprises an integrity signature.

6 . The method of claim 4 , wherein the extension is a non-standard extension in a data format only understandable by the data lakehouse.

7 . The method of claim 1 , further comprising:

receiving, by the one or more processors, a second request from the query engine or a second query engine to load a second data file to the data lakehouse;

retrieving, by the one or more processors, the second data file and information associated with the second data file;

generating, by the one or more processors, a second header/footer file containing the information associated with the second data file;

copying, by the one or more processors, the second data file to generate a copied data file; and

storing, by the one or more processors, the copied data file and the second header/footer file in the data lakehouse.

8 . The method of claim 1 , wherein generating the composite file and storing the composite file is performed automatically in response to a trigger mechanism.

9 . The method of claim 8 , wherein the trigger mechanism comprises at least one of a notification of a file addition or an instruction to list new files since a predetermined date.

10 . A system comprising:

one or more processors; and

one or more storage devices coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations for managing files in a data lakehouse, the operations comprising:

receiving a request from a query engine to load a data file to the data lakehouse;

retrieving the data file and information associated with the data file;

generating a header/footer file containing the information associated with the data file;

generating a composite file based on the data file and the header/footer file by concatenating the data file with the header/footer file, wherein at least one of a header or footer of the data file becomes available space in memory; and

storing the composite file in the data lakehouse.

11 . The system of claim 10 , wherein the information associated with the data file comprises at least one of a stable column identifier, a partition key, or an integrity signature.

12 . The system of claim 10 , wherein the header or footer of the data file becomes available space in memory in response to being concatenated with the header/footer file.

13 . The system of claim 10 , wherein the operations further comprise storing an extension in the available space in memory between the data file and the header/footer file of the composite file.

14 . The system of claim 13 , wherein the extension comprises an integrity signature.

15 . The system of claim 13 , wherein the extension is a non-standard extension in a data format only understandable by the data lakehouse.

16 . The system of claim 10 , wherein the operations further comprise:

receiving a second request from the query engine or a second query engine to load a second data file to the data lakehouse;

retrieving the second data file and information associated with the second data file;

generating a second header/footer file containing the information associated with the second data file;

copying the second data file to generate a copied data file; and

storing the copied data file and the second header/footer file in the data lakehouse.

17 . The system of claim 10 , wherein generating the composite file and storing the composite file is performed automatically in response to a trigger mechanism.

18 . A non-transitory computer readable medium for storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations for managing files in a data lakehouse, the operations comprising:

receiving a request from a query engine to load a data file to the data lakehouse;

retrieving the data file and information associated with the data file;

generating a header/footer file containing the information associated with the data file;

generating a composite file based on the data file and the header/footer file by concatenating the data file with the header/footer file, wherein at least one of a header or footer of the data file becomes available space in memory; and

storing the composite file in the data lakehouse.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2024
From: FANG, ZHOU; GUO, JIAN; HOTTELIER, THIBAUD; JOHNSON, ANOOP KOCHUMMEN; KORNFIELD, MICAH; LEVANDOSKI, JUSTIN; VOLOBUEV, YURI; ZHANG, YIWEI
To: GOOGLE LLC
Reel/Frame 068042/0001 →
References Cited (19)
US 11748318B1 · Cseri et al. · 2023 [cited by applicant]
US 12135621B1 · Thomas · 2024 [cited by examiner]
US 20200293193A1 · Littlefield et al. · 2020 [cited by applicant]
US 20210011891A1 · Soza · 2021 [cited by applicant]
US 20230140109A1 · Dasi et al. · 2023 [cited by applicant]
US 20230229658A1 · Newman · 2023 [cited by examiner]
US 20240330192A1 · Cardente · 2024 [cited by examiner]
N. William Rayner, Chapter 6—Data Handling, Editor(s): Eleftheria Zeggini, Andrew Morris, Analysis of Complex Disease Association Studies, Academic Press, 2011, pp. 87-94 (Year: 2011). [cited by examiner]
Rey et al., Seamless Integration of Parquet Files into Data Processing, Business, Technologie und Web (BTW 2023), Lecture Notes in Informatics (LNI), Gesellschaft für Informatik, Bonn 2023, pp. 235-258. (Year: 2023). [cited by examiner]
“Introducing Tabular” [online]. [Retrieved Apr. 26, 2024] Retrieved from the internet: <https://docs.tabular.io/en/introducing-tabular.html>. 4 pages. [cited by applicant]
“Migrating Hive workloads to Iceberg” Manual. Cloudera, Jan. 2023. 10 pages. [cited by applicant]
Armbrust, M., et al., “Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics”, Proceedings of CIDR. vol. 8., Jan. 2021. 8 pages. [cited by applicant]
Malone, J., “Iceberg Tables: Powering Open Standards with Snowflake Innovations” [online] Aug. 8, 2022. Retrieved from the internet: <https://www.snowflake.com/blog/iceberg-tables-powering-open-standards-with-snowflake-… [cited by applicant]
Strengholt, P., “Scalable Data Management with Microsoft Fabric and Microsoft Purview”, [online]. Apr. 8, 2024. Retrieved from the internet: <https://piethein.medium.com/scalable-data-management-with-microsoft-fabric-an… [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2025/024473 dated Jul. 15, 2025. 16 pages. [cited by applicant]
Levandoski et al. BigLake: BigQuery's Evolution toward a Multi-Cloud Lakehouse. Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1, ACMPUB27, New York, NY, USA, Jun. 9, 2024 (Jun. 9, 2024… [cited by applicant]
Mazumdar et al. The Data Lakehouse: Data Warehousing and More. arxiv.org, Cornell University Library, 201 OLIN Library Cornell University Ithaca, NY 14853, Oct. 12, 2023 (Oct. 12, 2023), 12 pages. [cited by applicant]
Tutorial: Loading and unloading Parquet data—Snowflake Documentation. Unload the table. Jun. 14, 2024 (Jun. 14, 2024), pp. 1-2, Retrieved from the Internet: <https://web.archive.org/web/20240614003323/https://docs.snowf… [cited by applicant]
Xinli Shang et al. Fast Copy-On-Write within Apache Parquet for Data Lakehouse ACID Upserts. Apr. 5, 2024 (Apr. 5, 2024), pp. 1-7, Retrieved from the Internet: <https://web.archive.org/web/20240405193844/https://www.ube… [cited by applicant]