IP Library Granted Patent US 10,339,016
Granted Patent B2
US 10,339,016 · App. 15/674,362 · Granted Jul 2, 2019

Chunk allocation

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,339,016
App. No.
15/674,362
Granted
Jul 2, 2019
Kind
B2
Abstract

Methods and systems for identifying a set of disks within a cluster and then storing a plurality of data chunks into the set of disks such that the placement of the plurality of data chunks within the cluster optimizes failure tolerance and storage system performance for the cluster are described. The plurality of data chunks may be generated using replication of data (e.g., n-way mirroring) or application of erasure coding to the data (e.g., using a Reed-Solomon code or a Low-Density Parity-Check code). The topology of the cluster including the physical arrangement of the nodes and disks within the cluster and status information for the nodes and disks within the cluster (e.g., information regarding disk fullness, disk performance, and disk age) may be used to identify the set of disks in which to store the plurality of data chunks.

Claims (57)

1. A method for operating a data management system, comprising:

generating a plurality of data chunks associated with a snapshot of a real or virtual machine;

identifying a set of preferred disks out of a plurality of disks within a cluster that stores other data chunks associated with the real or virtual machine;

acquiring disk status information for the plurality of disks within the cluster;

determining a plurality of failure domains for the plurality of disks using the disk status information;

identifying a set of disks out of the plurality of disks within the cluster in which to store the plurality of data chunks based on the set of preferred disks and the plurality of failure domains; and

writing the plurality of data chunks to the set of disks.

2. The method of claim 1 , wherein:

the disk status information includes disk ages for the plurality of disks; and

the determining the plurality of failure domains for the plurality of disks includes grouping the plurality of disks into the plurality of failure domains based on the disk ages of the plurality of disks.

3. The method of claim 1 , wherein:

the disk status information includes disk MTTF values for the plurality of disks; and

the determining the plurality of failure domains for the plurality of disks includes grouping the plurality of disks into the plurality of failure domains based on the disk MTTP values for the plurality of disks.

4. The method of claim 1 , wherein:

the identifying the set of disks includes generating a priority list of disks by acquiring a hierarchical resource pool, generating the plurality of failure domains using the hierarchical resource pool, and interleaving disks from the plurality of failure domains.

5. The method of claim 1 , wherein:

the identifying the set of disks includes identifying the set of disks that maximizes a total disk allocation function that weighs failure domain distances between the disks of the set of disks.

6. The method of claim 5 , wherein:

a first failure domain distance of the failure domain distances between a first disk of the set of disks and a second disk of the set of disks corresponds with a number of edges within a failure domain hierarchy separating a first disk-level failure domain of the plurality of failure domains that includes the first disk and a second disk-level failure domain of the plurality of failure domains that includes the second disk.

7. The method of claim 1 , wherein:

the disk status information includes disk fullness values for the plurality of disks; and

the identifying the set of disks includes identifying the set of disks based on the disk fullness values for the plurality of disks.

8. The method of claim 1 , wherein:

the identifying the set of preferred disks includes identifying the set of preferred disks that stores other data chunks associated with the snapshot.

9. The method of claim 1 , wherein:

the identifying the set of preferred disks includes identifying the set of preferred disks that stores other data chunks associated with a snapshot chain of the real or virtual machine.

10. The method of claim 1 , wherein:

the generating the plurality of data chunks includes acquiring the snapshot and applying erasure coding techniques to the snapshot.

11. The method of claim 1 , wherein:

the generating the plurality of data chunks includes partitioning the snapshot into segments and replicating the segments.

12. The method of claim 1 , wherein:

the writing the plurality of data chunks includes concurrently writing each data chunk of the plurality of data chunks into the set of disks.

13. The method of claim 1 , wherein:

the snapshot comprises a virtual machine snapshot.

14. The method of claim 1 , wherein:

the snapshot corresponds with a forward incremental for the real or virtual machine.

15. A data management system, comprising:

a memory configured to store a snapshot of a real or virtual machine; and

one or more processors configured to generate a plurality of data sets associated with the snapshot and identify a set of preferred disks out of a plurality of disks within a cluster that stores other data sets associated with the real or virtual machine, the one or more processors configured to acquire disk status information for the plurality of disks within the cluster and determine a plurality of failure domains for the plurality of disks based on the disk status information, the one or more processors configured to identify a set of disks out of the plurality of disks within the cluster in which to store the plurality of data sets based on the set of preferred disks and the plurality of failure domains, the one or more processors configured to cause the plurality of data sets to be concurrently written to the set of disks.

16. The data management system of claim 15 , wherein:

the disk status information includes disk ages for the plurality of disks; and

the one or more processors configured to group the plurality of disks into the plurality of failure domains based on the disk ages of the plurality of disks.

17. The data management system of claim 15 , wherein:

the disk status information includes disk MTTF values for the plurality of disks; and

the one or more processors configured to group the plurality of disks into the plurality of failure domains based on the disk MTTP values for the plurality of disks.

18. The data management system of claim 15 , wherein:

the one or more processors configured to identify the set of disks that maximizes a total disk allocation function that weighs failure domain distances between the disks of the set of disks, a first failure domain distance of the failure domain distances between a first disk of the set of disks and a second disk of the set of disks corresponds with a number of edges within a hierarchical resource pool separating a first disk-level failure domain that includes the first disk and a second disk-level failure domain that includes the second disk.

19. The data management system of claim 15 , wherein:

the disk status information includes disk fullness values for the plurality of disks; and

the one or more processors configured to identify the set of disks based on the disk fullness values for the plurality of disks.

20. One or more storage devices containing processor readable code for programming one or more processors to perform a method for operating a data management system, the processor readable code comprising:

processor readable code configured to acquire a plurality of data chunks associated with a snapshot of a virtual machine;

processor readable code configured to identify a set of preferred disks out of a plurality of disks within a cluster that stores other data chunks associated with the virtual machine;

processor readable code configured to acquire disk status information for the plurality of disks within the cluster, the disk status information includes disk ages for the plurality of disks;

processor readable code configured to determine a plurality of failure domains for the plurality of disks using the disk status information to group the plurality of disks into the plurality of failure domains based on the disk ages for the plurality of disks;

processor readable code configured to identify a set of disks out of the plurality of disks within the cluster in which to store the plurality of data chunks using the set of preferred disks and the plurality of failure domains; and

processor readable code configured to store the plurality of data chunks using the set of disks.

Assignments (3)
RELEASE OF SECURITY INTEREST IN PATENT COLLATERAL AT REEL/FRAME NO. 60333/0323 Recorded Jun 13, 2025
From: GOLDMAN SACHS BDC, INC., AS COLLATERAL AGENT
To: RUBRIK, INC.
Reel/Frame 071565/0602 →
GRANT OF SECURITY INTEREST IN PATENT RIGHTS Recorded Jun 10, 2022
From: RUBRIK, INC.
To: GOLDMAN SACHS BDC, INC., AS COLLATERAL AGENT
Reel/Frame 060333/0323 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 10, 2017
From: JUNIWAL, GARVIT; JAIN, GAURAV; GEE, ADAM
To: RUBRIK, INC.
Reel/Frame 043264/0497 →
Cited By (3)
US 12,271,269 US 12,298,941 US 12,650,951