IP Library › Granted Patent US 12,405,924
Granted Patent B2
US 12,405,924 · App. 18/389,337 · Granted Sep 2, 2025

Managed tables for data lakes

Inventors: Victor Sergeyevich Agababov (Seattle, WA); Shuang Guan (Sunnyvale, CA); Thibaud Hottelier (Seattle, WA); Anoop Kochummen Johnson (Fremont, CA); Justin Levandoski (Seattle, WA); Bigang Li (Redmond, WA); Yuri Volobuev (Walnut Creek, CA)
Assignee: Google LLC
G06F16/1805G06F12/0253G06F16/221G06F16/2358G06F16/2365G06F16/2379G06F16/283
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,405,924
App. No.
18/389,337
Granted
Sep 2, 2025
Kind
B2
Abstract

Aspects of the disclosure are directed to merging data lake openness with scalable metadata for managed tables in a cloud database platform, allowing for atomicity, consistency, isolation, and durability (ACID) transactions, performant data manipulation language (DML), higher throughput stream ingestion, data consistency, schema evolution, time travel, clustering, fine-grained security, and/or automatic storage optimization. Table data is stored in various open-source file formats in cloud storage while physical metadata of the table data is stored in a scalable metadata storage system.

Claims (49)

1. A method for processing queries, comprising:

receiving, by one or more processors, a request from a query engine to write one or more tuples;

writing, by the one or more processors, the one or more tuples to a write-optimized storage in a row-oriented format in a distributed file system that supports file appends;

converting, by the one or more processors, the one or more tuples to one or more data files in a columnar-oriented format compatible with the query engine;

storing, by the one or more processors, the one or more data files in a read-optimized cloud storage in the columnar-oriented format compatible with the query engine; and

committing, by the one or more processors, the write as an addition to a table transaction log stored in the distributed file system by writing a row-level addition to the table transaction log.

2. The method of claim 1 , further comprising performing, by the one or more processors, one or more maintenance tasks in the distributed file system.

3. The method of claim 2 , wherein the one or more maintenance tasks comprise garbage collection of one or more tuples stored in the write-optimized storage.

4. The method of claim 1 , further comprising:

converting, by one or more processors, the table transaction log to the columnar-oriented format compatible with the query engine; and

storing, by one or more processors, the table transaction log in the read-optimized cloud storage compatible with the query engine.

5. The method of claim 1 , wherein the query engine is a data lake query engine.

6. The method of claim 1 , wherein the committing of the write to the table transaction log occurs exactly once.

7. The method of claim 1 , further comprising:

receiving, by the one or more processors, a request from a query engine to read the one or more data files; and

reading, by the one or more processors, the one or more data files through an API that provides read after write semantics.

8. A system comprising:

one or more processors; and

one or more storage devices coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations for processing queries, the operations comprising:

receiving a request from a query engine to write one or more tuples;

writing tuples of the one or more tuples to a write-optimized storage in a row-oriented format in a distributed file system that supports file appends;

converting the one or more tuples to one or more data files in a columnar-oriented format compatible with the query engine;

storing the one or more data files in a read-optimized cloud storage in the columnar-oriented format compatible with the query engine; and

committing the write as an addition to a table transaction log stored in the distributed file system by writing a row-level addition to the table transaction log.

9. The system of claim 8 , wherein the operations further comprise performing one or more maintenance tasks in the distributed file system.

10. The system of claim 9 , wherein the one or more maintenance tasks comprise garbage collection of one or more tuples stored in the write-optimized storage.

11. The system of claim 8 , wherein the operations further comprise:

converting the table transaction log to the columnar-oriented format compatible with the query engine; and

storing the table transaction log in the read-optimized cloud storage compatible with the query engine.

12. The system of claim 8 , wherein the query engine is a data lake query engine.

13. The system of claim 8 , wherein the committing of the write to the table transaction log occurs exactly once.

14. The system of claim 8 , wherein the operations further comprise:

receiving a request from a query engine to read one or more additional data files; and

reading the one or more additional data files through an API that provides read after write semantics.

15. A non-transitory computer readable medium for storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations for processing queries, the operations comprising:

receiving a request from a query engine to write one or more tuples;

writing the one or more tuples to a write-optimized storage in a row-oriented format in a distributed file system that supports file appends;

converting the one or more tuples to one or more data files in a columnar-oriented format compatible with the query engine;

storing the one or more data files in a read-optimized cloud storage in the columnar-oriented format compatible with the query engine; and

committing the write as an addition to a table transaction log stored in the distributed file system by writing a row-level addition to the table transaction log.

16. The non-transitory computer readable medium of claim 15 , wherein the operations further comprise performing one or more maintenance tasks in the distributed file system.

17. The non-transitory computer readable medium of claim 16 , wherein the one or more maintenance tasks comprise garbage collection of one or more tuples stored in the write-optimized storage.

18. The non-transitory computer readable medium of claim 15 , wherein the operations further comprise:

converting the table transaction log to the columnar-oriented format compatible with the query engine; and

storing the table transaction log in the read-optimized cloud storage compatible with the query engine.

19. The non-transitory computer readable medium of claim 15 , wherein the committing of the write to the table transaction log occurs exactly once.

20. The non-transitory computer readable medium of claim 15 , wherein the operations further comprise:

receiving a request from a query engine to read one or more additional data files; and

reading the one or more additional data files through an API that provides read after write semantics.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: AGABABOV, VICTOR SERGEYEVICH; GUAN, SHUANG; HOTTELIER, THIBAUD; JOHNSON, ANOOP KOCHUMMEN; LEVANDOSKI, JUSTIN; LI, BIGANG; VOLOBUEV, YURI
To: GOOGLE LLC
Reel/Frame 065557/0177 →
Continuity (2)
Provisional Application 63535811 · Aug 31, 2023
Related Publication 20250077478A1 · Mar 6, 2025
References Cited (19)
US 7792822B2 · Galindo-Legaria et al. · 2010 [cited by applicant]
US 11106661B2 · Grabs · 2021 [cited by examiner]
US 11455290B1 · Brahmadesam · 2022 [cited by examiner]
US 11615083B1 · Attaluri · 2023 [cited by examiner]
US 11860869B1 · Hwang · 2024 [cited by examiner]
US 20140279838A1 · Tsirogiannis · 2014 [cited by examiner]
US 20160147859A1 · Lee · 2016 [cited by examiner]
US 20170116237A1 · Zhang · 2017 [cited by examiner]
US 20200364201A1 · Cseri · 2020 [cited by examiner]
US 20210342067A1 · Meister et al. · 2021 [cited by applicant]
US 20220083978A1 · Weindling · 2022 [cited by examiner]
US 20220327131A1 · Akidau et al. · 2022 [cited by applicant]
US 20220382674A1 · Wang et al. · 2022 [cited by applicant]
US 20230185688A1 · Edara et al. · 2023 [cited by applicant]
US 20230409545A1 · Gupta · 2023 [cited by examiner]
US 20240311350A1 · Pandya · 2024 [cited by examiner]
Edara et al. Big Metadata: When Metadata is Big Data. 2021. PVLDB, 14(12): pp. 3083-3095. [cited by applicant]
Behm et al. Photon: A Fast Query Engine for Lakehouse Systems. Proceedings of the 26th ACM International Conference on Hybrid Systems: Computation and Control, ACMPUB27, New York, NY, USA, Jun. 10, 2022 (Jun. 10, 2022),… [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2024/035961 dated Sep. 10, 2024. 18 pages. [cited by applicant]