IP Library Granted Patent US 12,602,500
Granted Patent B2
US 12,602,500 · App. 18/439,073 · Granted Apr 14, 2026

Encapsulating access algorithms for data processing engines

Inventors: Ronald J. Barber (San Jose, CA); Richard Sefton Sidle (Ottawa, CA); Viktor Giannakouris Salalidis (Ithaca, NY); Mir Hamid Pirahesh (San Jose, CA); Berthold Reinwald (San Jose, CA)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F21/6218G06F16/258
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,500
App. No.
18/439,073
Granted
Apr 14, 2026
Kind
B2
Abstract

A first data processing engine with an unpublished data format may store data on a shared storage system, where the data is in an unpublished data format. The first data processing engine may generate code, executable by a second data processing engine, where the code is configured to allow the second data processing engine to directly access a portion of the data without having access to knowledge of the unpublished data format.

Claims (43)

1 . A system comprising:

a memory storing program instructions; and

a processor in communication with the memory, the processor being configured to execute the program instructions to perform processes comprising:

storing, by a first data processing engine with an unpublished data format, data on a shared storage system; and

generating, by the first data processing engine, executable code, wherein the code is configured to allow a second data processing engine to directly access and read a portion of the data without having access to knowledge of the unpublished data format.

2 . The system of claim 1 , wherein the memory stores further program instructions, and wherein the processor is configured to execute the further program instructions to perform the processes further comprising:

generating worker node partitions to process portions of the generated code in parallel.

3 . The system of claim 1 , wherein the memory stores further program instructions, and wherein the processor is configured to execute the further program instructions to perform the processes further comprising:

receiving a SQL query from a client requesting access to the data.

4 . The system of claim 3 , wherein the memory stores additional program instructions, and wherein the processor is configured to execute the additional program instructions to perform the processes further comprising:

generating a subquery from the query prompting the first data processing engine to generate the code.

5 . The system of claim 1 , wherein the memory stores further program instructions, and wherein the processor is configured to execute the further program instructions to perform the processes further comprising:

processing the generated code in batches.

6 . The system of claim 1 , wherein the memory stores further program instructions, and wherein the processor is configured to execute the further program instructions to perform the processes further comprising:

bridging the generated code to be processed by the second data processing engine.

7 . The system of claim 1 , wherein the generated code comprises metadata with data size and data location.

8 . A method comprising:

storing, by a first data processing engine with an unpublished data format, data on a shared storage system; and

generating, by the first data processing engine, executable code, wherein the code is configured to allow a second data processing engine to directly access and read a portion of the data without having access to knowledge of the unpublished data format.

9 . The method of claim 8 , wherein the method further comprises:

generating worker node partitions to process portions of the generated code in parallel.

10 . The method of claim 8 , wherein the method further comprises:

receiving a SQL query from a client requesting access to the data.

11 . The method of claim 10 , wherein the method further comprises:

generating a subquery from the query prompting the first data processing engine to generate the code.

12 . The method of claim 8 , wherein the method further comprises:

processing the generated code in batches.

13 . The method of claim 8 , wherein the method further comprises:

bridging the generated code to be processed by the second data processing engine.

14 . The method of claim 8 , wherein the generated code contains metadata with data size and data location.

15 . A computer program product comprising one or more computer readable storage media having program instructions collectively embodied therewith, the program instructions executable by one or more processors to cause the one or more processors to perform a method, the method comprising:

storing, by a first data processing engine with an unpublished data format, data on a shared storage system; and

generating, by the first data processing engine, executable code, wherein the code is configured to allow a second data processing engine to directly access and read a portion of the data without having access to knowledge of the unpublished data format.

16 . The computer program product of claim 15 , further comprising additional program instructions collectively stored on the one or more computer readable storage media and configured to cause the one or more processors to perform the method further comprising:

generating worker node partitions to process portions of the generated code in parallel.

17 . The computer program product of claim 15 , further comprising additional program instructions stored on the one or more computer readable storage media and configured to cause the one or more processors to perform the method further comprising:

receiving a SQL query from a client requesting access to the data.

18 . The computer program product of claim 17 , further comprising additional program instructions stored on the one or more computer readable storage media and configured to cause the one or more processors to perform the method further comprising:

generating a subquery from the query prompting the first data processing engine to generate the code.

19 . The computer program product of claim 15 , further comprising additional program instructions stored on the one or more computer readable storage media and configured to cause the one or more processors to perform the method further comprising:

processing the generated code in batches.

20 . The computer program product of claim 15 , further comprising additional program instructions stored on the one or more computer readable storage media and configured to cause the one or more processors to perform the method further comprising:

bridging the generated code to be processed by the second data processing engine.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 13, 2024
From: BARBER, RONALD J.; SIDLE, RICHARD SEFTON; GIANNAKOURIS SALALIDIS, VIKTOR; PIRAHESH, MIR HAMID; REINWALD, BERTHOLD
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 066448/0373 →
Continuity (1)
Related Publication 20250258949A1 · Aug 14, 2025
References Cited (10)
US 8380738B2 · Tatemura · 2013 [cited by applicant]
US 10503923B1 · Gupta · 2019 [cited by examiner]
US 11449508B2 · Potharaju · 2022 [cited by applicant]
US 20070011134A1 · Langseth · 2007 [cited by examiner]
US 20190044953A1 · Holsinger · 2019 [cited by examiner]
Armbrust et al. “Lakehouse: a new generation of open platforms that unify data warehousing and advanced analytics.” Proceedings of CIDR. vol. 8. 2021,8 pages. [cited by applicant]
Nambiar et al., An Overview of Data Warehouse and Data Lake in Modern Enterprise Data Management. Big Data and Cognitive Computing 6.4, Nov. 7, 2022, 24 pages. [cited by applicant]
Orescanin et al., “Data lakehouse—a novel step in analytics architecture.” 2021 44th International Convention on Information, Communication and Electronic Technology (MIPRO). IEEE, 202, pp. 1242-1246. [cited by applicant]
Shankaran et al. “The Gluten Open-Source Software Project: Modernizing Java-based Query Engines for the Lakehouse Era.” 2023, 7 pages. [cited by applicant]
Simitsis et al., “The History, Present, and Future of ETL Technology.” [Test-of-Time Award—Invited Talk], 2023, 10 pages. [cited by applicant]