IP Library › Granted Patent US 11,288,132
Granted Patent B1
US 11,288,132 · App. 16/588,782 · Granted Mar 29, 2022

Distributing multiple phases of deduplication processing amongst a set of nodes within a clustered storage environment

Inventors: Abhishek Rajimwale (San Jose, CA); George Mathew (Belmont, CA)
Assignee: EMC IP Holding Company LLC
G06F11/1464G06F9/5083G06F11/0709G06F11/1453G06F11/1469
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,288,132
App. No.
16/588,782
Filed
Sep 30, 2019
Granted
Mar 29, 2022
Kind
B1
Art Unit
2136
USPC
711/162
Abstract

Described is a system for distributing multiple phases of a deduplication processing amongst of set of nodes. The system may perform a load-balancing in configurations where multiple generations of backup data are redirected to the same host node, and thus, require the host node to perform certain storage processes such as writing new backup data to its associated physical storage. Accordingly, the system may perform an initial (or first phase) processing on a first node that is selected based on resource usage or classification (e.g. metadata storing node). The system may then perform a subsequent (or second phase) processing on a second, or host node, that is selected based on the node already storing previous generations of the backup data. Accordingly, the system still redirects processing to a host node, but provides the ability to delegate certain deduplication operations to additional nodes.

Claims (53)

1. A system comprising:

one or more processors; and

a non-transitory computer readable medium storing a plurality of instructions, which when executed, cause the one or more processors to:

receive backup data to be stored on one or more nodes of a clustered storage system;

determine a first node of the clustered storage system already stores a first backup file associated with the received backup data, the first backup file being a previous generation of the received backup data;

select a second node of the clustered storage system based on a workload balancing factor or a policy associated with the second node;

initiate the selected second node of the clustered storage system to perform a first phase of a multiple phase deduplication processing of the received backup data the first phase deduplication processing including at least computing fingerprints of the received backup data;

select, after the first phase deduplication processing of the received backup data, the first node to perform a second phase of the multiple phase deduplication process based on the first node already storing the first backup file associated with the received backup data; and

initiate the selected first node to perform the second phase of the multiple phase deduplication processing by redirecting the multiple phase deduplication processing to the first node, the second phase deduplication processing including at least writing the backup data for storage as a second backup file on the first node.

2. The system of claim 1 , wherein the plurality of instructions, when executed, further cause the one or more processors to:

determine a resource usage associated with at least some of the nodes; and

select the second node to perform the first phase deduplication processing based on the resource usage.

3. The system of claim 1 , wherein the plurality of instructions, when executed, further cause the one or more processors to:

determine the second node is assigned as a metadata node for the clustered storage system; and

select the second node to perform the first phase deduplication processing based on the second node being assigned as the metadata node for the clustered storage system.

4. The system of claim 3 , wherein the metadata node stores file system information of the backup files stored by clustered storage system.

5. The system of claim 4 , wherein determining the first node of the clustered storage system already stores the first backup file includes accessing the file system information stored by the metadata node.

6. The system of claim 1 , wherein determining the first node of the clustered storage system already stores the first backup file includes matching a first data source identifier associated with the first backup file with a second data source identifier provided with the received backup data.

7. The system of claim 1 , wherein the second phase deduplication processing further includes filtering the backup data by identifying duplicate data segments already stored by the clustered storage system.

8. The system of claim 1 , wherein the plurality of instructions, when executed, further cause the one or more processors to:

select the second node to perform the first phase deduplication processing based on the second node being associated with a data source of the backup data.

9. A method comprising:

receiving, by a system, backup data to be stored on one or more nodes of a clustered storage system;

determining, by the system, a first node of the clustered storage system already stores a first backup file associated with the received backup data, the first backup file being a previous generation of the received backup data;

selecting, by the system, a second node of the clustered storage system based on a workload balancing factor or policy associated with the second node;

initiate the selected second node of the clustered storage system to perform a first phase of a multiple phase deduplication processing of the received backup data the first phase deduplication processing including at least computing fingerprints of the received backup data;

selecting, by the system, after the first phase deduplication processing of the received backup data, the first node to perform a second phase of the multiple phase deduplication process based on the first node already storing the first backup file associated with the received backup data; and

initiating, by the system, the selected first node to perform the second phase of the multiple phase deduplication processing by redirecting the multiple phase deduplication processing to the first node, the second phase deduplication processing including at least writing the backup data for storage as a second backup file on the first node.

10. The method of claim 9 , further comprising:

determining a resource usage associated with at least some of the nodes; and

selecting the second node to perform the first phase deduplication processing based on the resource usage.

11. The method of claim 9 , further comprising:

determining the second node is assigned as a metadata node for the clustered storage system; and

selecting the second node to perform the first phase deduplication processing based on the second node being assigned as the metadata node for the clustered storage system.

12. The method of claim 11 , wherein the metadata node stores file system information of the backup files stored by the clustered storage system, and wherein determining the first node of the clustered storage system already stores the first backup file includes accessing the file system information stored by the metadata node.

13. The method of claim 9 , wherein determining the first node of the clustered storage system already stores the first backup file includes matching a first data source identifier associated with the first backup file with a second data source identifier provided with the received backup data.

14. The method of claim 9 , wherein the second phase deduplication processing further includes filtering the backup data by identifying duplicate data segments already stored by the clustered storage system.

15. A computer program product comprising a non-transitory computer-readable medium having a computer-readable program code embodied therein to be executed by one or more processors, the program code including instructions to:

receive backup data to be stored on one or more nodes of a clustered storage system;

determine a first node of the clustered storage system already stores a first backup file associated with the received backup data, the first backup file being a previous generation of the received backup data;

select a second node of the clustered storage system based on a workload balancing factor or policy associated with the second node;

initiate the selected second node of the clustered storage system to perform a first phase of a multiple phase deduplication processing of the received backup data the first phase deduplication processing including at least computing fingerprints of the received backup data;

select, after the first phase deduplication processing of the received backup data, the first node to perform a second phase of the multiple phase deduplication process based on the first node already storing the first backup file associated with the received backup data; and

initiate the selected first node to perform the second phase of the multiple phase deduplication processing by redirecting the multiple phase deduplication processing to the first node, the second phase deduplication processing including at least writing the backup data for storage as a second backup file on the first node.

16. The computer program product of claim 15 , wherein the program code includes further instructions to:

determine a resource usage associated with at least some of the nodes; and

select the second node to perform the first phase deduplication processing based on the resource usage.

17. The computer program product of claim 15 , wherein the program code includes further instructions to:

determine the second node is assigned as a metadata node for the clustered storage system; and

select the second node to perform the first phase deduplication processing based on the second node being assigned as the metadata node for the clustered storage system.

18. The computer program product of claim 17 , wherein the metadata node stores file system information of the backup files stored by the clustered storage system, and wherein determining the first node of the clustered storage system already stores the first backup file includes accessing the file system information stored by the metadata node.

19. The computer program product of claim 15 , wherein determining the first node of the clustered storage system already stores the first backup file includes matching a first data source identifier associated with the first backup file with a second data source identifier provided with the received backup data.

20. The computer program product of claim 15 , wherein the second phase deduplication processing further includes filtering the backup data by identifying duplicate data segments already stored by the clustered storage system.

Assignments (9)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053546/0001) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC IP HOLDING COMPANY LLC
Reel/Frame 071642/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (051302/0528) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO WYSE TECHNOLOGY L.L.C.); SECUREWORKS CORP.
Reel/Frame 060438/0593 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053311/0169) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
Reel/Frame 060438/0742 →
RELEASE OF SECURITY INTEREST AT REEL 051449 FRAME 0728 Recorded Nov 2, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.; SECUREWORKS CORP.; EMC CORPORATION
Reel/Frame 058002/0010 →
SECURITY INTEREST Recorded Jun 5, 2020
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS COLLATERAL AGENT
Reel/Frame 053311/0169 →
SECURITY AGREEMENT Recorded Apr 22, 2020
From: CREDANT TECHNOLOGIES INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 053546/0001 →
SECURITY AGREEMENT Recorded Dec 31, 2019
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.; SECUREWORKS CORP.; EMC CORPORATION
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 051449/0728 →
PATENT SECURITY AGREEMENT (NOTES) Recorded Dec 16, 2019
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.; SECUREWORKS CORP.
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS COLLATERAL AGENT
Reel/Frame 051302/0528 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2019
From: RAJIMWALE, ABHISHEK; MATHEW, GEORGE
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 050570/0485 →
Cited By (2)
US 12,216,550 US 12,430,304