IP Library Granted Patent US 11,340,839
Granted Patent B2
US 11,340,839 · App. 17/336,628 · Granted May 24, 2022

Sub-cluster recovery using a partition group index

Inventors: Rohit Shekhar (Sunnyvale, CA); Hyo Jun Kim (San Jose, CA); Prasenjit Sarkar (Los Gatos, CA); Maohua Lu (Fremont, CA); Ajaykrishna Raghavan (Santa Clara, CA); Pin Zhou (San Jose, CA)
Assignee: Rubrik, Inc.
G06F3/067G06F3/0604G06F3/0641G06F11/1469G06F16/00G06F11/1456G06F2201/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,340,839
App. No.
17/336,628
Granted
May 24, 2022
Kind
B2
Abstract

The method disclosed includes scanning data items stored in the first plurality of nodes of a first cluster. While scanning, creating a partition group index indexing the data items into a plurality of partition groups. Each partition group corresponds to a node of the first plurality of nodes and comprises a subset of data items stored in the node. Storing the index. Instantiating a second cluster, comprising generating per node data, for each node of a second plurality of nodes, based on mappings between the partition groups and the first plurality of nodes. Identifying the data items included in the partition groups according to the partition group index and loading the data items included in the partition groups onto the second plurality of nodes.

Claims (57)

1. A method for sub-cluster recovery in a data storage environment having a first plurality of nodes, the method comprising:

scanning data items stored in the first plurality of nodes of a first cluster;

while scanning, creating a partition group index indexing the data items into a plurality of partition groups, each partition group corresponds to a node of the first plurality of nodes and comprises a subset of data items stored in the node;

storing the index;

instantiating a second cluster comprising generating per node data, for each node of a second plurality of nodes, based on mappings between the partition groups and the first plurality of nodes;

identifying the data items included in the partition groups according to the partition group index, and

loading the data items included in the partition groups onto the second plurality of nodes.

2. The method of claim 1 , further comprising recovering each node of the first plurality of nodes by identifying the data items stored in the node based on the mappings between the plurality of partition groups and the first plurality of nodes.

3. The method of claim 2 , further comprising recovering each node of the first plurality of nodes by restoring the identified data items onto the node.

4. The method of claim 1 , wherein the instantiating the second cluster comprises generating a partition group to node mapping for the second cluster.

5. The method of claim 1 , further comprising:

while scanning the data items, identifying duplicate data items in the cluster;

deduplicating the duplicate data items;

repackaging each of the duplicate data items into respective deduplicated data units; and

storing the deduplicated data units to a secondary data repository.

6. The method of claim 5 , further comprising:

determining a degree of duplicates of the data items in the cluster; and

comparing the degree of duplicates to a predetermined level of consistency,

wherein the deduplication is performed responsive to determining the degree of duplicates is greater than the predetermined level of consistency.

7. The method of claim 5 , wherein the partition groups include deduplicated data items.

8. The method of claim 5 , wherein storing the deduplicated data units comprises:

storing a data version of the data items and;

compiling the deduplicated data units into the data version of the data items.

9. The method of claim 1 , wherein the data items are stored in a No SQL data store.

10. A system comprising:

at least one processor and executable instructions accessible on a computer-readable medium that, when executed, cause the at least one processor to perform operations comprising:

scanning data items stored in a first plurality of nodes of a first cluster;

while scanning, creating a partition group index indexing the data items into a plurality of partition groups, each partition group corresponds to a node of the first plurality of nodes and comprises a subset of data items stored in the node;

storing the index;

instantiating a second cluster comprising generating per node data, for each node of a second plurality of nodes, based on mappings between the partition groups and the first plurality of nodes;

identifying the data items included in the partition groups according to the partition group index, and

loading the data items included in the partition groups onto the second plurality of nodes.

11. The system of claim 10 , further comprising recovering each node of the first plurality of nodes by identifying the data items stored in the node based on the mappings between the plurality of partition groups and the first plurality of nodes.

12. The system of claim 11 , further comprising recovering each node of the first plurality of nodes by restoring the identified data items onto the node.

13. The system of claim 10 , wherein the instantiating the second cluster comprises generating a partition group to node mapping for the second cluster.

14. The system of claim 10 , further comprising:

while scanning the data items, identifying duplicate data items in the cluster;

deduplicating the duplicate data items;

repackaging each of the duplicate data items into respective deduplicated data units; and

storing the deduplicated data units to a secondary data repository.

15. The system of claim 14 , further comprising:

determining a degree of duplicates of the data items in the cluster; and

comparing the degree of duplicates to a predetermined level of consistency,

wherein the deduplication is performed responsive to determining the degree of duplicates is greater than the predetermined level of consistency.

16. The system of claim 14 , wherein the partition groups include deduplicated data items.

17. The system of claim 14 , wherein storing the deduplicated data units comprises:

storing a data version of the data items and;

compiling the deduplicated data units into the data version of the data items.

18. The system of claim 10 , wherein the data items are stored in a No SQL data store.

19. A non-transitory machine-readable medium storing a set of instructions that, when executed by a processor, causes a machine to perform operations comprising:

scanning data items stored in a first plurality of nodes of a first cluster;

while scanning, creating a partition group index indexing the data items into a plurality of partition groups, each partition group corresponds to a node of the first plurality of nodes and comprises a subset of data items stored in the node;

storing the index;

instantiating a second cluster comprising generating per node data, for each node of a second plurality of nodes, based on mappings between the partition groups and the first plurality of nodes;

identifying the data items included in the partition groups according to the partition group index, and

loading the data items included in the partition groups onto the second plurality of nodes.

20. The non-transitory machine-readable medium of claim 19 , further comprising recovering each of the first plurality of nodes by identifying the data items stored in the node based on the mappings between the partition groups and the first plurality of nodes.

Assignments (4)
RELEASE OF SECURITY INTEREST IN PATENT COLLATERAL AT REEL/FRAME NO. 60333/0323 Recorded Jun 13, 2025
From: GOLDMAN SACHS BDC, INC., AS COLLATERAL AGENT
To: RUBRIK, INC.
Reel/Frame 071565/0602 →
GRANT OF SECURITY INTEREST IN PATENT RIGHTS Recorded Jun 10, 2022
From: RUBRIK, INC.
To: GOLDMAN SACHS BDC, INC., AS COLLATERAL AGENT
Reel/Frame 060333/0323 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2021
From: SHEKHAR, ROHIT; KIM, HYO JUN; SARKAR, PRASENJIT; LU, MAOHUA; RAGHAVAN, AJAYKRISHNA; ZHOU, PIN
To: DATOS IO INC.
Reel/Frame 056439/0145 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2021
From: DATOS IO INC.
To: RUBRIK, INC.
Reel/Frame 056439/0229 →