IP Library Granted Patent US 11,080,207
Granted Patent B2
US 11,080,207 · App. 15/616,186 · Granted Aug 3, 2021

Caching framework for big-data engines in the cloud

Inventors: Joydeep Sen Sarma (Bangalore, IN); Rajat Venkatesh (Bengaluru, IN); Shubham Tagra (Bangalore, IN)
Assignee: Qubole, Inc.
G06F12/128G06F12/0813G06F12/0868G06F12/0871G06F2212/1032G06F2212/1041G06F2212/466G06F2212/604G06F2212/621
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,080,207
App. No.
15/616,186
Granted
Aug 3, 2021
Kind
B2
Abstract

The present invention is generally directed to a caching framework that provides a common abstraction across one or more big data engines, comprising a cache filesystem including a cache filesystem interface used by applications to access cloud storage through a cache subsystem, the cache filesystem interface in communication with a big data engine extension and a cache manager; the big data engine extension, providing cluster information to the cache filesystem and working with the cache filesystem interface to determine which nodes cache which part of a file; and a cache manager for maintaining metadata about the cache, the metadata comprising the status of blocks for each file. The invention may provide common abstraction across big data engines that does not require changes to the setup of infrastructure or user workloads, allows sharing of cached data and caching only the parts of files that are required, can process columnar format.

Claims (13)

1. A caching framework for storing metadata independent of a specific big data engine, the caching framework providing a common abstraction to share cached data across multiple big data engines, comprising a cache filesystem stored in a non-transitory computer readable storage medium comprising:

a cache filesystem interface, used by applications to access cloud storage through a cache subsystem, the cache filesystem interface in communication with a big data engine extension and a cache manager;

the big data engine extension, providing cluster information to the cache filesystem and working with the cache filesystem interface to determine which nodes cache which part of a file; and

a cache manager, responsible for maintaining metadata about the cache, the metadata comprising the status of blocks for each file and stored separately from the block and configured to ensure that only one copy of a block is written to the cache.

2. The caching framework of claim 1 , wherein determining which nodes cache with part of a file by the cache filesystem is performed using consistent hashing to reduce the impact of rebalance the cache when a node joins or leaves the cluster.

3. The caching framework of claim 1 , wherein the cache manager executes cache eviction.

4. The caching framework of claim 3 , wherein cache eviction is configured based upon eviction policies, eviction policies selected from the group consisting of Least Recently Used First, or Least Frequently Used First, time-based eviction, and disk usage-based eviction.

5. The caching framework of claim 1 , wherein the cache manager is further configured to pin certain files or parts of files to cache so that such files or parts of files cannot be evicted.

6. The caching framework of claim 5 , wherein the cache manager comprises pinning policies provided by a user or a plugin which are used, at least in part, to determine which parts of a file are important and pin such parts to the cache.

7. The caching framework of claim 6 , wherein the parts of a file that are important comprise headers and footers containing metadata.

8. The caching framework of claim 1 , wherein the cache filesystem communicates with the cache manager to retrieve and update block metadata.

9. The caching framework of claim 8 , wherein the cache manager ensures that only one copy of the block is written to the cache in a concurrent environment.

10. The caching framework of claim 1 wherein the multiple big data engines comprise at least two engines are selected from the group consisting of Map-Reduce, Spark, Presto, Hive, and Tez.

Assignments (6)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE CONVEYANCE TO READ: RELEASE OF SECOND LIEN SECURITY INTEREST IN SPECIFIED PATENTS RECORDED AT RF 054498/0130 PREVIOUSLY RECORDED ON REEL 70689 FRAME 837. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded Apr 2, 2025
From: JEFFERIES FINANCE LLC, AS COLLATERAL AGENT
To: QUBOLE INC.
Reel/Frame 070706/0009 →
RELEASE OF FIRST LIEN SECURITY INTEREST IN SPECIFIED PATENTS RECORDED AT RF 054498/0115 Recorded Mar 31, 2025
From: JEFFERIES FINANCE LLC, AS COLLATERAL AGENT
To: QUBOLE INC.
Reel/Frame 070689/0831 →
RELEASE OF FIRST LIEN SECURITY INTEREST IN SPECIFIED PATENTS RECORDED AT RF 054498/0130 Recorded Mar 31, 2025
From: JEFFERIES FINANCE LLC, AS COLLATERAL AGENT
To: QUBOLE INC.
Reel/Frame 070689/0837 →
FIRST LIEN SECURITY AGREEMENT Recorded Nov 23, 2020
From: QUBOLE INC.
To: JEFFERIES FINANCE LLC
Reel/Frame 054498/0115 →
SECOND LIEN SECURITY AGREEMENT Recorded Nov 23, 2020
From: QUBOLE INC.
To: JEFFERIES FINANCE LLC
Reel/Frame 054498/0130 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2017
From: SARMA, JOYDEEP SEN; VENKATESH, RAJAT; TAGRA, SHUBHAM
To: QUBOLE INC.
Reel/Frame 042940/0684 →
Continuity (2)
Provisional Application 62346627 · Jun 7, 2016
Related Publication 20170351620A1 · Dec 7, 2017