IP Library Granted Patent US 11,675,743
Granted Patent B2
US 11,675,743 · App. 17/590,469 · Granted Jun 13, 2023

Web-scale distributed deduplication

Inventor: Hariprasad Bhasker Rao Mankude (San Ramon, CA)
Assignee: Cohesity, Inc.
G06F16/1752G06F16/182
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,675,743
App. No.
17/590,469
Granted
Jun 13, 2023
Kind
B2
Abstract

Approaches for parallelized data deduplication. An instruction to perform data deduplication on a plurality of files is received. The plurality of files is organized into two or more work sets that each correspond to a subset of the plurality of files. Responsibility for performing each of said two or more work sets is assigned to a set of nodes in a cluster of nodes. The nodes may be physical nodes or virtual nodes. Each node in the set performs data deduplication on a different work set. In performing data deduplication, each node may store metadata describing where shared chunks of data are maintained in a distributed file system. The shared chunks of data are two or more sequences of bytes which appear in two or more of said plurality of files.

Claims (49)

1. A system, comprising:

a deduplication coordinator;

a first cluster node coupled to the deduplication coordinator, wherein the first cluster node includes a plurality of compute containers; and

one or more other cluster nodes coupled to the deduplication coordinator, wherein the one or more other cluster nodes include one or more corresponding compute containers,

a parallel database to store metadata in a chunk identifier table which is accessible from all nodes of a cluster, wherein the chunk identifier table includes a list of fingerprints,

wherein:

each compute container of the plurality compute containers and the one or more corresponding compute containers is configured to:

perform, in parallel, deduplication with respect to a corresponding assigned subset of files at least in part by:

create, using a fingerprinting algorithm, variable sized chunks of data associated with a file of the corresponding assigned subset of files and identify boundaries associated with the variable sized chunks;

create fingerprints of the variable sized chunks using a hash algorithm;

determine whether a fingerprint of the fingerprints already exists or is present in the parallel database; and

in response to a determination that the fingerprint does not already exist or is not present in the parallel database, update the chunk identifier table with information that enables the data chunk associated with the fingerprint to be located, wherein the fingerprint is associated with a file offset and length information; and

generate corresponding deduplication statistics associated with the corresponding assigned subset of files; and

the deduplication coordinator is configured to aggregate from each compute container the corresponding deduplication statistics.

2. The system of claim 1 , wherein the corresponding assigned subset of files is assigned to a cluster node based on available bandwidth or processing power.

3. The system of claim 1 , wherein the first cluster node is configured to assign a subset of files to one of the plurality of compute containers in response to a determination that a specific compute container on the first cluster node is not assigned the subset of files.

4. The system of claim 1 , wherein each compute container is configured to transmit the corresponding deduplication statistics.

5. The system of claim 1 , wherein each compute container is configured to scan the corresponding assigned subset of files.

6. The system of claim 1 , wherein a chunking algorithm is applied to a stream of the data associated with the file to identify the boundaries associated with the variable sized chunks.

7. The system of claim 1 , wherein the variable sized chunks of data are compressed.

8. The system of claim 7 , wherein the compressed variable sized chunks of data are written to a distributed file system.

9. The system of claim 1 , wherein the corresponding assigned subset of files is associated with a directory or folder.

10. The system of claim 1 , wherein the first cluster node is assigned an additional subset of files after each of the other cluster nodes is assigned an initial subset of files.

11. The system of claim 1 , wherein the deduplication coordinator includes a user interface.

12. The system of claim 11 , wherein the user interface is configured to receive a specification of files to which the deduplication is to be performed.

13. The system of claim 12 , wherein the specification indirectly specifies files to which the deduplication is to be performed.

14. The system of claim 12 , wherein the specification directly specifies files to which the deduplication is to be performed.

15. The system of claim 1 , wherein the parallel database includes a plurality of tables that includes information that indicates whether a file has been deduplicated.

16. The system of claim 1 , wherein the parallel database includes a plurality of tables that includes information about how to reconstruct a file if the file has been deduplicated.

17. The system of claim 1 , wherein the parallel database includes a plurality of tables that include a global table having the fingerprint as a row key.

18. The system of claim 1 , wherein the deduplication coordinator assigns a particular subset of files to a first compute container having a cumulative default size.

19. A method, comprising:

assigning a corresponding subset of files to a plurality of cluster nodes, wherein a first cluster node of the plurality of cluster nodes includes a plurality of computer containers and one or more other cluster nodes of the plurality of cluster nodes include one or more corresponding compute containers;

performing, in parallel by each compute container of the plurality of compute containers and the one or more corresponding compute containers, deduplication with respect to a corresponding assigned subset of files at least in part by:

creating, using a fingerprinting algorithm, variable sized chunks of data associated with a file of the corresponding assigned subset of files and identifying boundaries associated with the variable sized chunks;

creating fingerprints of the variable sized chunks using a hash algorithm;

determining whether a fingerprint of the fingerprints already exists or is present in a parallel database;

in response to a determination that the fingerprint does not already exist or is not present in the parallel database, updating a chunk identifier table with information that enables the data chunk associated with the fingerprint to be located, wherein the chunk identifier table is included in the parallel database that is accessible from all of the plurality of cluster nodes of a cluster and stores metadata, wherein the fingerprint is associated with a file offset and length information; and

generating corresponding deduplication statistics associated with the corresponding assigned subset of files; and

aggregating from each of the compute containers the corresponding deduplication statistics.

20. A computer program product embodied in a non-transitory computer readable medium and comprising computer instructions for:

assigning a corresponding subset of files to a plurality of cluster nodes, wherein a first cluster node of the plurality of cluster nodes includes a plurality of computer containers and one or more other cluster nodes of the plurality of cluster nodes include one or more corresponding compute containers;

performing, in parallel by each compute container of the plurality of compute containers and the one or more corresponding compute containers, deduplication with respect to a corresponding assigned subset of files at least in part by:

creating, using a fingerprinting algorithm, variable sized chunks of data associated with a file of the corresponding assigned subset of files and identify boundaries associated with the variable sized chunks;

creating fingerprints of the variable sized chunks using a hash algorithm;

determining whether a fingerprint of the fingerprints already exists or is present in a parallel database;

in response to a determination that the fingerprint does not already exist or is not present in the parallel database, updating a chunk identifier table with information that enables the data chunk associated with the fingerprint to be located, wherein the chunk identifier table is included in the parallel database that is accessible from all of the plurality of cluster nodes of a cluster and stores metadata, wherein the fingerprint is associated with a file offset and length information; and

generating corresponding deduplication statistics associated with the corresponding assigned subset of files; and

aggregating from each of the compute containers the corresponding deduplication statistics.

Assignments (6)
TERMINATION AND RELEASE OF INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Dec 10, 2024
From: FIRST-CITIZENS BANK & TRUST COMPANY (AS SUCCESSOR TO SILICON VALLEY BANK)
To: COHESITY, INC.
Reel/Frame 069584/0498 →
SECURITY INTEREST Recorded Dec 9, 2024
From: VERITAS TECHNOLOGIES LLC; COHESITY, INC.
To: JPMORGAN CHASE BANK. N.A.
Reel/Frame 069890/0001 →
SECURITY INTEREST Recorded Sep 23, 2022
From: COHESITY, INC.
To: SILICON VALLEY BANK, AS ADMINISTRATIVE AGENT
Reel/Frame 061509/0818 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 6, 2022
From: IMANIS DATA INC.
To: COHESITY, INC.
Reel/Frame 059863/0868 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 6, 2022
From: MANKUDE, HARIPRASAD BHASKER RAO
To: TALENA, INC.
Reel/Frame 059908/0656 →
CHANGE OF NAME Recorded May 6, 2022
From: TALENA, INC.
To: IMANIS DATA INC.
Reel/Frame 059908/0661 →
Continuity (4)
Continuation 16705089 · Dec 5, 2019
Continuation 14876579 · Oct 6, 2015
Provisional Application 62060367 · Oct 6, 2014
Related Publication 20220261379A1 · Aug 18, 2022