IP Library Granted Patent US 9,898,478
Granted Patent B2
US 9,898,478 · App. 14/673,586 · Granted Feb 20, 2018

Distributed deduplicated storage system

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,898,478
App. No.
14/673,586
Granted
Feb 20, 2018
Kind
B2
Abstract

A distributed, deduplicated storage system according to certain embodiments is arranged in a parallel configuration including multiple deduplication nodes. Deduplicated data is distributed across the deduplication nodes. The deduplication nodes can be networked together and communicate with one another according using a light-weight, customized communication scheme (e.g., a scheme based on FTP or HTTP). In some cases, deduplication management information including deduplication signatures and/or other metadata is stored separately from the deduplicated data in deduplication management nodes, improving performance and scalability.

Claims (34)

1. A method of performing a storage operation in a distributed deduplicated storage system, comprising:

creating with a first deduplication node of a plurality of deduplication nodes, a first hash signature of a first data block of a plurality of data blocks associated with a file, and a first header that at least identifies a first media agent that stored a copy of the first data block;

creating with a second deduplication node of the plurality of deduplication nodes, a second hash signature of at least a second data block associated with the file, and a second header that at least identifies a second media agent that stored a copy of the second data block;

storing at the first deduplication node a copy of the second hash signature, and a copy of the second header created by the second deduplication node;

receiving a first request from a client computing device to restore the file comprising the plurality of data blocks;

determining with the first deduplication node that the copy of the second data block was stored by the second media agent based at least in part on accessing the copy of the second hash signature, and the copy of the second header stored in association with the first deduplication node; and

sending a second request for the second data block to the second media agent that requests the second data block, wherein the second request comprises at least the second header.

2. The method of claim 1 further comprising determining which of the plurality of deduplication nodes to send a request comprises performing a modulo operation on the second hash signature of the second data block.

3. The method claim 1 wherein the second deduplication node uses the second hash signature to locate a storage location of the second data block.

4. The method of claim 1 wherein the first deduplication node further provides block offset information associated with the second data block to the second media agent.

5. The method of claim 1 wherein the first deduplication node further provides a path identifier associated with the second data block to the second media agent.

6. The method of claim 1 wherein the second media agent transmits a copy of the particular data block in response to said request for the second data block, and wherein the method further comprises receiving the transmitted copy of the second data block from the second media agent.

7. The method of claim 1 wherein the first and second media agents communicate hash signatures and headers without using network shares.

8. The method of claim 3 wherein the first and second media agents communicate hash signatures and headers and links between one another without having a shared static mount path configuration.

9. The method of claim 1 wherein sending the second request for the second data block is performed using a file-transfer protocol (FTP)-based service routine.

10. The method of claim 1 wherein sending the second request for the second data block is performed using a hyper-text transfer protocol (HTTP)-based service routine.

11. A distributed deduplicated storage system, comprising:

a plurality of deduplication nodes each comprising one or more processors and storage, the deduplication nodes in communication with one another,

a first deduplication node of the plurality of deduplication nodes creates a first hash signature of a first data block of a plurality of data blocks associated with a file and a first header that at least identifies a first media agent that stored a copy of the first data block; and

a second deduplication node of the plurality of deduplication nodes creates a second hash signature of at least a second data block associated with the file and a second header that at least identifies a second media agent that stored a copy of the second data block,

wherein the first deduplication node stores a copy of second hash signature and the second header created by the second depulication node;

computer hardware configured to:

receive a request for the file comprised of a plurality of data blocks;

determine with the first deduplication node that the copy of the second data block was stored by the second media agent based at least in part on accessing the copy of the second hash signature and the copy of the second header stored in association with the first deduplication node; and

sending a second request for the second data block to the second media agent wherein the second request comprises at least the second header.

12. The distributed deduplicated storage system of claim 11 further comprising determining which of the plurality of deduplication nodes to send a request comprises performing a modulo operation on the second hash signature of the second data block.

13. The distributed deduplicated storage system of claim 11 wherein the second deduplication node uses the second hash signature to locate a storage location of the second data block.

14. The distributed deduplicated storage system of claim 11 wherein the first deduplication node further provides block offset information associated with the second data block to the second media agent.

15. The distributed deduplicated storage system of claim 11 wherein the first deduplication node further provides a path identifier associated with the second data block to the second media agent.

16. The distributed deduplicated storage system of claim 11 wherein the second media agent server transmits a copy of the second data block to the first media agent in response to the second request for the second data block.

17. The distributed deduplicated storage system of claim 11 wherein the first and second media agents communicate without using network shares.

18. The distributed deduplicated storage system of claim 11 wherein the first and second media agents communicate without having a shared static mount path configuration.

19. The distributed deduplicated storage system of claim 11 wherein the second media agent is configured to respond to the second request using a file-transfer protocol (FTP)-based service routine.

20. The distributed deduplicated storage system of claim 11 wherein the media agent is configured to perform the request using a hyper-text transfer protocol (HTTP)-based service routine.

Assignments (3)
SUPPLEMENTAL CONFIRMATORY GRANT OF SECURITY INTEREST IN UNITED STATES PATENTS Recorded Apr 16, 2025
From: COMMVAULT SYSTEMS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 070864/0344 →
SECURITY INTEREST Recorded Dec 13, 2021
From: COMMVAULT SYSTEMS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 058496/0836 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2015
From: VIJAYAN, MANOJ KUMAR; KOTTOMTHARAYIL, RAJIV; ATTARDE, DEEPAK RAGHUNATH
To: COMMVAULT SYSTEMS, INC.
Reel/Frame 035845/0902 →