IP Library Granted Patent US 9,678,973
Granted Patent B2
US 9,678,973 · App. 14/515,020 · Granted Jun 13, 2017

Multi-node hybrid deduplication

Inventors: Ronald Ray Trimble (Acton, MA); Jeffrey V. Tofano (San Jose, CA); Thomas R. Ramsdell (San Jose, CA); Jon Christopher Kennedy (Marlborough, MA)
Assignee: HITACHI DATA SYSTEMS CORPORATION
G06F17/30156
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,678,973
App. No.
14/515,020
Granted
Jun 13, 2017
Kind
B2
Abstract

According to at least one embodiment, a data storage system is provided. The data storage system includes memory, at least one processor in data communication with the memory, and a deduplication director component executable by the at least one processor. The deduplication director component is configured to receive data for storage on the data storage system, analyze the data to determine whether the data is suitable for at least one of summary-based deduplication, content-based deduplication, and no deduplication, and store, in a common object store, at least one of the data and a reference to duplicate data stored in the common object store.

Claims (50)

1. A data storage system comprising:

memory;

at least one processor in data communication with the memory; and

a deduplication director component executable by the at least one processor and configured to:

receive data for storage on the data storage system;

analyze the data to determine whether or not the data is suitable for deduplication;

when it is determined that the data is suitable for deduplication, determine which deduplication process is suitable for the data from at least one of summary-based deduplication and content-based deduplication; and

store, in a common object store, at least one of the data and a reference to duplicate data stored in the common object store.

2. The data storage system of claim 1 , wherein the deduplication director component is configured to store, in the common object store, the data in response to determining that the data is not suitable for deduplication.

3. The data storage system of claim 1 , wherein the deduplication director component is configured to store, in the common object store, the reference to the duplicate data at least in part by executing a summary-based deduplication component in response to determining that the data is suitable for summary-based deduplication, the summary-based deduplication component being configured to:

generate a summary of the data;

identify the data as being a copy of the duplicate data at least in part by comparing the summary to a duplicate summary of the duplicate data; and

store, in the common object store, the reference to the duplicate data in response to identifying the data as being a copy of the duplicate data.

4. The data storage system of claim 1 , wherein the deduplication director component is configured to store, in the common object store, the reference to the duplicate data at least in part by executing a content-based deduplication component in response to determining that the data is suitable for content-based deduplication, the content-based deduplication component being configured to:

identify the data as being a copy of the duplicate data by comparing the data to the duplicate data; and

store, in the common object store, the reference to the duplicate data in response to identifying the data as being a copy of the duplicate data.

5. The data storage system of claim 1 , wherein the deduplication director component is configured to analyze the data at least in part to identify at least one deduplication template to apply to the data, the at least one deduplication template specifying one or more deduplication options for the data.

6. The data storage system of claim 5 , wherein the at least one deduplication template specifies that at least one of a summary-based deduplication component and a content-based deduplication component processes the data.

7. The data storage system of claim 5 , wherein the deduplication director component is configured to analyze characteristics of the data at least in part to identify the at least one deduplication template to apply to the data, the characteristics of the data include at least one of file metadata and a source of the data.

8. The data storage system of claim 1 , further comprising a content-based deduplication component and a summary-based deduplication component, wherein the common object store stores a map that associates the data with one or more objects stored in the common object store and at least one of the content-based deduplication component and the summary-based deduplication component is configured to deduplicate the map.

9. The data storage system of claim 8 , wherein the content-based deduplication component is configured to deduplicate the map at least in part by re-segmenting at least one data object associated with the map.

10. The data storage system of claim 1 , further comprising:

a file system interface component configured to receive one or more files for storage on the data storage system; and

a common object store component executable by the at least one processor and configured to store each of the one or more files as one or more data objects structured differently from the one or more files.

11. The data storage system of claim 10 , wherein the file system interface component includes a proxy file system interface configured to receive source level deduplication information.

12. The data storage system of claim 11 , wherein the proxy file system interface is further configured to determine whether subject data identified within the source level deduplication information is stored in the data storage system.

13. The data storage system of claim 1 , wherein the deduplication director component is further configured to determine whether deduplication should be performed as an inline process or as a post-process.

14. A method of deduplicating data using a data storage system, the method being executed on a hardware device and comprising:

receiving, by the data storage system, data for storage on the data storage system;

analyzing the data to determine whether or not the data is suitable for deduplication;

when it is determined that the data is suitable for deduplication, determine which deduplication process is suitable for the data from at least one of summary-based deduplication and content-based deduplication; and

storing, in a common object store, at least one of the data and a reference to duplicate data stored in the common object store.

15. The method of claim 14 , wherein analyzing the data includes identifying at least one deduplication template to apply to the data, the at least one deduplication template specifying one or more deduplication options for the data.

16. The method of claim 15 , wherein identifying the at least one deduplication template includes identifying at least one deduplication template that specifies that at least one of a summary-based deduplication component and a content-based deduplication component processes the data.

17. The method of claim 15 , wherein analyzing the data further includes analyzing characteristics of the data at least in part to identify the at least one deduplication template to apply to the data, the characteristics of the data include at least one of file metadata and a source of the data.

18. The method of claim 14 , wherein the common object store stores a map that associates the data with one or more objects stored in the common object store and the method further comprises deduplicating the map.

19. The method of claim 18 , wherein deduplicating the map includes re-segmenting at least one data object associated with the map.

20. The method of claim 14 , further comprising:

receiving one or more files for storage on the data storage system; and

storing each of the one or more files as one or more data objects structured differently from the one or more files.

21. The method of claim 20 , further comprising receiving source level deduplication information.

22. The method of claim 21 , further comprising determining whether subject data identified within the source level deduplication information is stored in the data storage system.

23. The method of claim 14 , wherein analyzing the data further includes determining whether deduplication should be performed as an inline process or as a post-process.

24. A non-transitory computer readable storage medium storing computer executable instructions configured to instruct at least one processor to execute a hybrid deduplication process, the computer executable instructions including instructions to:

receive data for storage on the data storage system;

analyze the data to determine whether or not the data is suitable deduplication;

when it is determined that the data is suitable for deduplication, determine which deduplication process is suitable for the data from at least one of summary-based deduplication and content-based deduplication; and

store, in a common object store, at least one of the data and a reference to duplicate data stored in the common object store.

25. The non-transitory storage medium of claim 24 , wherein the computer executable instructions further include instructions to determine whether deduplication should be performed as an inline process or as a post-process.

26. The non-transitory storage medium of claim 24 , wherein the computer executable instructions further include instructions to analyze characteristics of the data at least in part to identify at least one deduplication template to apply to the data, the characteristics of the data include at least one of file metadata and a source of the data, the at least one deduplication template specifying that at least one of a summary-based deduplication component and a content-based deduplication component processes the data.

Assignments (4)
MERGER Recorded Jan 28, 2020
From: HITACHI VANTARA CORPORATION
To: HITACHI VANTARA LLC
Reel/Frame 051719/0202 →
CHANGE OF NAME Recorded Feb 20, 2018
From: HITACHI DATA SYSTEMS CORPORATION
To: HITACHI VANTARA CORPORATION
Reel/Frame 045369/0785 →
MERGER Recorded Feb 9, 2017
From: SEPATON, INC.
To: HITACHI DATA SYSTEMS CORPORATION
Reel/Frame 041670/0829 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2015
From: TRIMBLE, RONALD RAY; TOFANO, JEFFREY V.; RAMSDELL, THOMAS R.; KENNEDY, JON CHRISTOPHER
To: SEPATON, INC.
Reel/Frame 035776/0087 →
Continuity (2)
Provisional Application 61891042 · Oct 15, 2013
Related Publication 20150106345A1 · Apr 16, 2015