IP Library › Granted Patent US 10,031,937
Granted Patent B2
US 10,031,937 · App. 14/952,307 · Granted Jul 24, 2018

Similarity based data deduplication of initial snapshots of data sets

Inventor: Lior Aronovich (Thornhill, CA)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F17/30371G06F3/0641G06F17/30088G06F17/30159
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,031,937
App. No.
14/952,307
Filed
Nov 25, 2015
Granted
Jul 24, 2018
Kind
B2
Art Unit
2159
USPC
707/692
Abstract

Embodiments for data deduplication of an initial snapshot of a data set in a storage system by a processor. An intra-snapshot similarity index, inclusive of representations of the data inside the initial snapshot, is built. The intra-snapshot similarity index is used for deduplication of the initial snapshot. The intra-snapshot similarity index is merged with a global similarity index.

Claims (49)

1. A method for data deduplication of an initial snapshot of a data set in a storage system by a processor, comprising:

building an intra-snapshot similarity index, inclusive only of representations of the data inside the initial snapshot;

using the intra-snapshot similarity index for deduplication of the initial snapshot in a chain of a plurality of snapshots by first using only the representations of the data within the intra-snapshot similarity index of the initial snapshot to perform the deduplication of the initial snapshot prior to using a global similarity index to perform the deduplication; wherein the global similarity index is used to perform the deduplication of the initial snapshot subsequent to using the intra-snapshot similarity index when a deduplication threshold is not met using the intra-snapshot similarity index; and

merging the intra-snapshot similarity index with the global similarity index by performing each of:

structurally merging the intra-snapshot index into the global similarity index,

bulk inserting entries of the intra-snapshot index into the global similarity index when unable to structurally merge the intra-snapshot index into the global similarity index, and

performing the merging of the intra-snapshot index with the global similarity index when deduplication processing of the initial snapshot is complete.

2. The method of claim 1 , further including:

for an input similarity unit, searching the intra-snapshot similarity index for similar data, and

deduplicating the input similarity unit with found data.

3. The method of claim 1 , wherein the intra-snapshot similarity index is built using a resolution that is higher than the resolution of the global similarity index.

4. The method of claim 3 , wherein sub-similarity units used to build and to search within the intra-snapshot similarity index are smaller than similarity units used for the global similarity index.

5. The method of claim 4 , further including searching high resolution representative values in the intra-snapshot similarity index, and identifying similar sub-units, for matching digests of an input similarity unit and digests of found sub-units to find identical data sections.

6. The method of claim 4 , further including calculating a representative value for an input similarity unit based on high resolution representative values of sub-units, the representative value searched in the global similarity index, and a corresponding similarity unit identified for matching digests of the input similarity unit and digests of a found similarity unit to find identical data sections.

7. The method of claim 1 , further including configuring the initial snapshot to not have a preceding snapshot of the same data set.

8. The method of claim 1 , further including configuring the intra-snapshot similarity index to reside in memory.

9. A system for data deduplication of an initial snapshot of a data set in a storage system, comprising:

a processor, operable in the storage system, wherein the processor:

builds an intra-snapshot similarity index, inclusive only of representations of the data inside the initial snapshot,

uses the intra-snapshot similarity index for deduplication of the initial snapshot in a chain of a plurality of snapshots by first using only the representations of the data within the intra-snapshot similarity index of the initial snapshot to perform the deduplication of the initial snapshot prior to using a global similarity index to perform the deduplication; wherein the global similarity index is used to perform the deduplication of the initial snapshot subsequent to using the intra-snapshot similarity index when a deduplication threshold is not met using the intra-snapshot similarity index, and

merges the intra-snapshot similarity index with the global similarity index by performing each of:

structurally merging the intra-snapshot index into the global similarity index,

bulk inserting entries of the intra-snapshot index into the global similarity index when unable to structurally merge the intra-snapshot index into the global similarity index, and

performing the merging of the intra-snapshot index with the global similarity index when deduplication processing of the initial snapshot is complete.

10. The system of claim 9 , wherein the processor:

for an input similarity unit, searches the intra-snapshot similarity index for similar data, and

deduplicates the input similarity unit with found data.

11. The system of claim 9 , wherein the intra-snapshot similarity index is built using a resolution that is higher than the resolution of the global similarity index.

12. The system of claim 11 , wherein sub-similarity units used to build and to search within the intra-snapshot similarity index are smaller than similarity units used for the global similarity index.

13. The system of claim 12 , wherein the processor searches high resolution representative values in the intra-snapshot similarity index, and identifies similar sub-units, for matching digests of an input similarity unit and digests of found sub-units to find identical data sections.

14. The system of claim 12 , wherein the processor calculates a representative value for an input similarity unit based on high resolution representative values of sub-units, the representative value searched in the global similarity index, and a corresponding similarity unit identified for matching digests of the input similarity unit and digests of a found similarity unit to find identical data sections.

15. The system of claim 9 , wherein the initial snapshot does not have a preceding snapshot of the same data set.

16. The system of claim 9 , wherein the intra-snapshot similarity index resides in memory.

17. A computer program product for data deduplication of an initial snapshot of a data set in a storage system by a processor, the computer program product comprising a computer-readable storage medium having computer-readable program code portions stored therein, the computer-readable program code portions comprising:

an executable portion that builds an intra-snapshot similarity index, inclusive only of representations of the data inside the initial snapshot;

an executable portion that uses the intra-snapshot similarity index for deduplication of the initial snapshot in a chain of a plurality of snapshots by first using only the representations of the data within the intra-snapshot similarity index of the initial snapshot to perform the deduplication of the initial snapshot prior to using a global similarity index to perform the deduplication; wherein the global similarity index is used to perform the deduplication of the initial snapshot subsequent to using the intra-snapshot similarity index when a deduplication threshold is not met using the intra-snapshot similarity index; and

an executable portion that merges the intra-snapshot similarity index with the global similarity index by performing each of:

structurally merging the intra-snapshot index into the global similarity index,

bulk inserting entries of the intra-snapshot index into the global similarity index when unable to structurally merge the intra-snapshot index into the global similarity index, and

performing the merging of the intra-snapshot index with the global similarity index when deduplication processing of the initial snapshot is complete.

18. The computer program product of claim 17 , further including an executable portion that:

for an input similarity unit, searches the intra-snapshot similarity index for similar data, and

deduplicates the input similarity unit with found data.

19. The computer program product of claim 17 , wherein the intra-snapshot similarity index is built using a resolution that is higher than the resolution of the global similarity index.

20. The computer program product of claim 19 , wherein sub-similarity units used to build and to search within the intra-snapshot similarity index are smaller than similarity units used for the global similarity index.

21. The computer program product of claim 20 , further including an executable portion that searches high resolution representative values in the intra-snapshot similarity index, and identifies similar sub-units, for matching digests of an input similarity unit and digests of found sub-units to find identical data sections.

22. The computer program product of claim 20 , further including an executable portion that calculates a representative value for an input similarity unit based on high resolution representative values of sub-units, the representative value searched in the global similarity index, and a corresponding similarity unit identified for matching digests of the input similarity unit and digests of a found similarity unit to find identical data sections.

23. The computer program product of claim 17 , wherein the initial snapshot does not have a preceding snapshot of the same data set.

24. The computer program product of claim 17 , wherein the intra-snapshot similarity index resides in memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 25, 2015
From: ARONOVICH, LIOR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 037142/0538 →
Continuity (1)
Related Publication 20170147648A1 · May 25, 2017