IP Library Granted Patent US 8,612,699
Granted Patent B2
US 8,612,699 · App. 12/823,255 · Granted Dec 17, 2013

Deduplication in a hybrid storage environment

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,612,699
App. No.
12/823,255
Granted
Dec 17, 2013
Kind
B2
Abstract

Deduplication in a hybrid storage environment includes determining characteristics of a first data set. The first data set is identified as redundant to a second data set and the second data set is stored in a first storage system. The deduplication also includes mapping the characteristics of the first data set to storage preferences, the storage preferences specifying storage system selections for storing data sets based upon attributes of the respective storage systems. The deduplication further includes storing, as a persistent data set, one of the first data set and the second data set in one of the storage systems identified from the mapping.

Claims (65)

1. A method for deduplication in a hybrid storage environment comprising disk and redundant array identification disk (RAID) storage systems, the method comprising:

determining characteristics of a first data set via a computer processor, the first data set is identified as redundant to a second data set, the second data set stored in a first storage system, and the first data set is not stored;

mapping the characteristics of the first data set to storage preferences, the storage preferences specifying selections of storage systems that include the first storage system and a second storage system, the storage systems configured for storing data sets based upon attributes of the storage systems; and

storing, as a persistent data set, one of the first data set and the second data set in one of the storage systems identified from the mapping, the storing including:

upon determining that the first storage system embodies a preferred disk technology from results of the mapping, maintaining storage of the second data set in the first storage system as the persistent data set and ignoring the first data set; and

upon determining that a second storage system of the storage systems embodies the preferred disk technology from results of the mapping, storing the first data set in the second storage system as the persistent data set and clearing the second data set from the first storage system.

2. The method of claim 1 , further comprising creating a logical reference pointer to the persistent data set.

3. The method of claim 1 , wherein determining the characteristics of the first data set is performed in response to a request to store the first data set as part of an inline deduplication process.

4. The method of claim 1 , wherein determining the characteristics of the first data set further includes accessing data recorded in at least one of an audit log and an access log associated with the hybrid storage environment.

5. The method of claim 1 , wherein preferred disk technologies include hard disk drives and solid state disk drives.

6. The method of claim 1 , wherein the characteristics include at least one of:

frequently read;

frequently written;

frequently read and written;

files accessed sequentially;

files access randomly; and

file size.

7. The method of claim 1 , wherein the RAID storage system includes RAID-configured logical units (LUNs), the method further comprising:

identifying a preferred RAID-configured LUN for storing as a persistent data set, one of the first data set and the second data set, the preferred RAID-configured LUN identified from results of the mapping.

8. The method of claim 7 , further comprising rebuilding lost data for a failed RAID configuration and performing an inline deduplication process on rebuilt data, the deduplication process comprising:

modifying an amount of resources allocated to the inline deduplication process as a function of a duration of the deduplication process in view of input/output performance with respect to a specified SLA ratio.

9. A system for deduplication in a hybrid storage environment comprising disk and redundant array identification disk (RAID) storage systems, the system comprising:

a computer processor; and

a deduplication application executing on the computer processor, the deduplication application implementing a method, comprising:

determining characteristics of a first data set, the first data set is identified as redundant to a second data set, the second data set stored in a first storage system, and the first data set is not stored;

mapping the characteristics of the first data set to storage preferences, the storage preferences specifying selections of storage systems that include the first storage system and a second storage system, the storage systems configured for storing data sets based upon attributes of the storage systems; and

storing, as a persistent data set, one of the first data set and the second data set in one of the storage systems identified from the mapping, the storing including:

upon determining that the first storage system embodies a preferred disk technology from results of the mapping, maintaining storage of the second data set in the first storage system as the persistent data set and ignoring the first data set; and

upon determining that a second storage system of the storage systems embodies the preferred disk technology from results of the mapping, storing the first data set in the second storage system as the persistent data set and clearing the second data set from the first storage system.

10. The system of claim 9 , wherein the method further comprises creating a logical reference pointer to the persistent data set.

11. The system of claim 9 , wherein determining the characteristics of the first data set is performed in response to a request to store the first data set as part of an inline deduplication process.

12. The system of claim 9 , wherein determining the characteristics of the first data set further includes accessing data recorded in at least one of an audit log and an access log associated with the hybrid storage environment.

13. The system of claim 9 , wherein preferred disk technologies include hard disk drives and solid state disk drives; and

wherein the characteristics include at least one of:

frequently read;

frequently written;

frequently read and written;

files accessed sequentially;

files access randomly; and

file size.

14. The system of claim 9 , wherein the RAID storage system includes RAID-configured logical units (LUNs), the method further comprising:

identifying a preferred RAID-configured LUN for storing as a persistent data set, one of the first data set and the second data set, the preferred RAID-configured LUN identified from results of the mapping; and

rebuilding lost data for a failed RAID configuration and performing an inline deduplication process on rebuilt data, the deduplication process comprising:

modifying an amount of resources allocated to the inline deduplication process as a function of a duration of the deduplication process in view of input/output performance with respect to a specified SLA ratio.

15. A computer program product for deduplication in a hybrid storage environment, the computer program product comprising a non-transitory storage medium encoded with machine-readable computer program code, which when executed by a computer, cause the computer to implement a method, the method comprising:

determining characteristics of a first data set, the first data set is identified as redundant to a second data set, the second data set stored in a first storage system, and the first data set is not stored;

mapping the characteristics of the first data set to storage preferences, the storage preferences specifying selections of storage systems that include the first storage system and a second storage system, the storage systems configured for storing data sets based upon attributes of the storage systems; and

storing, as a persistent data set, one of the first data set and the second data set in one of the storage systems identified from the mapping, the storing including:

upon determining that the first storage system embodies a preferred disk technology from results of the mapping, maintaining storage of the second data set in the first storage system as the persistent data set and ignoring the first data set; and

upon determining that a second storage system of the storage systems embodies the preferred disk technology from results of the mapping, storing the first data set in the second storage system as the persistent data set and clearing the second data set from the first storage system.

16. The computer program product of claim 15 , wherein the method further comprises creating a logical reference pointer to the persistent data set.

17. The computer program product of claim 15 , wherein determining the characteristics of the first data set is performed in response to a request to store the first data set as part of an inline deduplication process.

18. The computer program product of claim 15 , wherein determining the characteristics of the first data set further includes accessing data recorded in at least one of an audit log and an access log associated with the hybrid storage environment.

19. The computer program product of claim 15 , wherein preferred disk technologies include hard disk drives and solid state disk drives.

20. The computer program product of claim 15 , wherein the characteristics include at least one of:

frequently read;

frequently written;

frequently read and written;

files accessed sequentially;

files access randomly; and

file size.

21. The computer program product of claim 15 , wherein the RAID storage system includes RAID-configured logical units (LUNs), the method further comprising:

identifying a preferred RAID-configured LUN for storing as a persistent data set, one of the first data set and the second data set, the preferred RAID-configured LUN identified from results of the mapping.

22. The computer program product of claim 21 , further comprising rebuilding lost data for a failed RAID configuration and performing an inline deduplication process on rebuilt data, the deduplication process comprising:

modifying an amount of resources allocated to the inline deduplication process as a function of a duration of the deduplication process in view of input/output performance with respect to a specified SLA ratio.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2024
From: BEIJING PIANRUOJINGHONG TECHNOLOGY CO., LTD.
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 066565/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2023
From: AWEMANE LTD.
To: BEIJING PIANRUOJINGHONG TECHNOLOGY CO., LTD.
Reel/Frame 064501/0498 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2021
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: AWEMANE LTD.
Reel/Frame 057991/0960 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 25, 2010
From: JAIN, BHUSHAN P; MUSIAL, JOHN G.; NAGPAL, ABHINAY R.; PATIL, SANDEEP R.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 024593/0573 →