IP Library Granted Patent US 11,340,838
Granted Patent B2
US 11,340,838 · App. 17/336,535 · Granted May 24, 2022

Sub-cluster recovery using a partition group index

Inventors: Rohit Shekhar (Sunnyvale, CA); Hyo Jun Kim (San Jose, CA); Prasenjit Sarkar (Los Gatos, CA); Maohua Lu (Fremont, CA); Ajaykrishna Raghavan (Santa Clara, CA); Pin Zhou (San Jose, CA)
Assignee: Rubrik, Inc.
G06F3/067G06F3/0604G06F3/0641G06F11/1469G06F16/00G06F11/1456G06F2201/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,340,838
App. No.
17/336,535
Granted
May 24, 2022
Kind
B2
Abstract

The method disclosed is for instantiating a second cluster based on a first cluster. For at least one node of a second plurality of nodes, generating per node data based on mappings between a plurality of partition groups and a first plurality of nodes, the first plurality of nodes corresponding to the first cluster. The method further discloses identifying data items included in the plurality of partition groups based on the mappings between the plurality of partition groups and the first plurality of nodes. The method further discloses each partition group corresponding to a node of the first plurality of nodes and comprising a subset of data items stored in the node. The method further discloses loading the data items included in the plurality of partition groups onto the second plurality of nodes, the second plurality of nodes corresponding to the second cluster.

Claims (51)

1. A method comprising:

instantiating a second cluster based on a first cluster, the instantiating the second cluster comprising:

for at least one node of a second plurality of nodes, generating per node data based on mappings between a plurality of partition groups and a first plurality of nodes, the first plurality of nodes corresponding to the first cluster;

identifying data items included in the plurality of partition groups based on the mappings between the plurality of partition groups and the first plurality of nodes, each partition group corresponding to a node of the first plurality of nodes and comprising a subset of data items stored in the node; and

loading the data items included in the plurality of partition groups onto the second plurality of nodes, the second plurality of nodes corresponding to the second cluster.

2. The method of claim 1 , further comprising recovering each node of the first plurality of nodes by identifying the data items stored in the node based on the mappings between the plurality of partition groups and the first plurality of nodes.

3. The method of claim 2 , further comprising recovering each node of the first plurality of nodes by restoring the identified data items onto the node.

4. The method of claim 1 , wherein the instantiating the second cluster comprises generating a partition group to node mapping for the second cluster.

5. The method of claim 1 , further comprising:

while scanning the data items, identifying duplicate data items in the cluster;

deduplicating the duplicate data items;

repackaging each of the duplicate data items into respective deduplicated data units; and

storing the deduplicated data units to a secondary data repository.

6. The method of claim 5 , further comprising:

determining a degree of duplicates of the data items in the cluster; and

comparing the degree of duplicates to a predetermined level of consistency,

wherein the deduplication is performed responsive to determining the degree of duplicates is greater than the predetermined level of consistency.

7. The method of claim 5 , wherein the partition groups include deduplicated data items.

8. The method of claim 5 , wherein storing the deduplicated data units comprises:

storing a data version of the data items and;

compiling the deduplicated data units into the data version of the data items.

9. The method of claim 1 , wherein the data items are stored in a No SQL data store.

10. A system comprising:

at least one processor and executable instructions accessible on a computer- readable medium that, when executed, cause the at least one processor to perform operations comprising:

instantiating a second cluster based on a first cluster, the instantiating the second cluster comprising:

for at least one node of a second plurality of nodes, generating per node data based on mappings between a plurality of partition groups and a first plurality of nodes, the first plurality of nodes corresponding to the first cluster;

identifying data items included in the plurality of partition groups based on the mappings between the plurality of partition groups and the first plurality of nodes, each partition group corresponding to a node of the first plurality of nodes and comprising a subset of data items stored in the node; and

loading the data items included in the plurality of partition groups onto the second plurality of nodes, the second plurality of nodes corresponding to the second cluster.

11. The system of claim 10 , further comprising recovering each node of the first plurality of nodes by identifying the data items stored in the node based on the mappings between the plurality of partition groups and the first plurality of nodes.

12. The system of claim 11 , further comprising recovering each node of the first plurality of nodes by restoring the identified data items onto the node.

13. The system of claim 10 , wherein the instantiating the second cluster comprises generating a partition group to node mapping for the second cluster.

14. The system of claim 10 , further comprising:

while scanning the data items, identifying duplicate data items in the cluster;

deduplicating the duplicate data items;

repackaging each of the duplicate data items into respective deduplicated data units; and

storing the deduplicated data units to a secondary data repository.

15. The system of claim 14 , further comprising:

determining a degree of duplicates of the data items in the cluster; and

comparing the degree of duplicates to a predetermined level of consistency,

wherein the deduplication is performed responsive to determining the degree of duplicates is greater than the predetermined level of consistency.

16. The system of claim 14 , wherein the partition groups include deduplicated data items.

17. The system of claim 14 , wherein storing the deduplicated data units comprises:

storing a data version of the data items and;

compiling the deduplicated data units into the data version of the data items.

18. The system of claim 10 , wherein the data items are stored in a No SQL data store.

19. A non-transitory machine-readable medium storing a set of instructions that, when executed by a processor, causes a machine to perform operations comprising:

instantiating a second cluster based on a first cluster, the instantiating the second cluster comprising:

for at least one node of a second plurality of nodes, generating per node data based on mappings between a plurality of partition groups and a first plurality of nodes, the first plurality of nodes corresponding to the first cluster;

identifying data items included in the plurality of partition groups based on the mappings between the plurality of partition groups and the first plurality of nodes, each partition group corresponding to a node of the first plurality of nodes and comprising a subset of data items stored in the node; and

loading the data items included in the plurality of partition groups onto the second plurality of nodes, the second plurality of nodes corresponding to the second cluster.

20. The non-transitory machine-readable medium of claim 19 , further comprising recovering each of the first plurality of nodes by identifying the data items stored in the node based on the mappings between the partition groups and the first plurality of nodes.

Assignments (4)
RELEASE OF SECURITY INTEREST IN PATENT COLLATERAL AT REEL/FRAME NO. 60333/0323 Recorded Jun 13, 2025
From: GOLDMAN SACHS BDC, INC., AS COLLATERAL AGENT
To: RUBRIK, INC.
Reel/Frame 071565/0602 →
GRANT OF SECURITY INTEREST IN PATENT RIGHTS Recorded Jun 10, 2022
From: RUBRIK, INC.
To: GOLDMAN SACHS BDC, INC., AS COLLATERAL AGENT
Reel/Frame 060333/0323 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2021
From: SHEKHAR, ROHIT; KIM, HYO JUN; SARKAR, PRASENJIT; LU, MAOHUA; RAGHAVAN, AJAYKRISHNA; ZHOU, PIN
To: DATOS IO INC.
Reel/Frame 056438/0947 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2021
From: DATOS IO INC.
To: RUBRIK, INC.
Reel/Frame 056439/0038 →