IP Library Granted Patent US 9,189,421
Granted Patent B2
US 9,189,421 · App. 13/610,567 · Granted Nov 17, 2015

System and method for implementing a hierarchical data storage system

Inventors: Richard Testardi (Boulder, CO); Maurilio Cometto (Palo Alto, CA); Kuriakose George Kulangara (Pune, IN)
Assignee: STORSIMPLE, INC.
G06F12/122G06F3/065G06F3/067G06F3/0608G06F3/0641G06F11/1448G06F11/1453G06F11/1456G06F11/1464G06F12/0808G06F2201/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,189,421
App. No.
13/610,567
Granted
Nov 17, 2015
Kind
B2
Abstract

A system and method for efficiently storing data both on-site and off-site in a cloud storage system. Data read and write requests are received by a cloud data storage system. The cloud storage system has at least three data storage layers. A first high-speed layer, a second efficient storage layer, and a third off-site storage layer. The first high-speed layer stores data in raw data blocks. The second efficient storage layer divides data blocks from the first layer into data slices and eliminates duplicate data slices. The third layer stores data slices at an off-site location.

Claims (37)

1. A data storage system, comprising:

a first data storage layer, the first data storage layer comprising data that can be accessed at a high data rate with a low latency;

a second data storage layer, said second data storage layer comprising de-duplicated data evicted from the first data storage layer, wherein the second storage layer has a higher retrieval latency than the first data storage layer;

a barrier storage area, the barrier storage area comprising a subset of the de-duplicated data that is stored for a settlement period, wherein the settlement period represents a time after which the de-duplicated data is made accessible; and

a third data storage layer coupled to the barrier storage area and configured to receive the subset of the de-duplicated data, said third data storage layer having a higher retrieval latency than said second data storage layer;

wherein a first data item is stored in said first, second, or third data storage layer based upon frequency of use of said first data item.

2. The data storage system of claim 1 , wherein said third data storage layer comprises a remote cloud storage service.

3. The data storage system of claim 1 , wherein said data storage system further comprises:

a first background process associated with said first data storage layer, said first background process evicting data from said first data storage layer when an amount of free space in said first storage layer is less than a first threshold level.

4. The data storage system for of claim 3 , wherein said data storage system further comprises:

a second background process associated with said second data storage layer, said second background process evicting data from said second data storage layer when an amount of free storage space in said second storage layer is less than a second threshold level.

5. The data storage system of claim 3 , wherein said first background process evicts data from said first storage layer using a least recently used heuristic.

6. The data storage system of claim 3 , wherein said first background process evicts data from said first storage layer using a least recently allocated heuristic.

7. The data storage system of claim 1 , wherein the de-duplicated data comprises data slices that are subsets of said data blocks with duplicate data slices eliminated.

8. The data storage system of claim 7 , wherein the subset of the de-duplicated data comprises data slices from said second data layer that have been compressed.

9. The data storage system of claim 1 , wherein a second data item may be marked as high-priority such that said second data item is kept in said first data storage layer, regardless of the amount of free space in said first storage layer.

10. The data storage system of claim 8 , wherein said third data storage layer comprises a remote cloud storage system.

11. The data storage system of claim 7 , wherein said data slices are each assigned a statistically unique identifier.

12. A system, comprising:

a first data storage layer in a memory, the first data storage layer comprising data blocks;

a second data storage layer in the memory, the second data storage layer comprising de-duplicated data evicted from the first data storage layer, wherein the second storage layer has a higher retrieval latency than the first data storage layer; and

a barrier storage area in the memory, the barrier storage area comprising a subset of the de-duplicated data that is stored for a settlement period before being deleted, wherein the subset of the de-duplicated data is selected for transfer to a third storage layer operated by a cloud storage provider, wherein the third data storage layer has a higher retrieval latency than the second data storage layer.

13. The system of claim 12 , further comprising:

a third data storage layer coupled to the barrier storage area and configured to receive the subset of the de-duplicated data, said third storage layer having a higher retrieval latency and data compression ratio than said second data storage layer;

wherein a first data item is stored in said first, second, or third data storage layer based upon frequency of use of said first data item.

14. The system of claim 13 , wherein the settlement period is configurable dependent upon a cloud storage provider being used, wherein the settlement period represents a time after which the de-duplicated data stored by the cloud storage provider is made accessible, and wherein a given data item is stored in the first, second, or third data storage layer based upon frequency of use of the given data item.

15. The system of claim 12 , further comprising:

a first background process associated with said first data storage layer, said first background process evicting data from said first data storage layer when an amount of free space in said first storage layer is less than a first threshold level.

16. The system of claim 15 , further comprising:

a second background process associated with said second data storage layer, said second background process evicting data from said second data storage layer when an amount of free storage space in said second storage layer is less than a second threshold level.

17. The system of claim 15 , wherein said first background process evicts data from said first storage layer using a least recently used heuristic.

18. The system of claim 15 , wherein said first background process evicts data from said first storage layer using a least recently allocated heuristic.

19. The system of claim 18 , wherein the subset of the de-duplicated said third data format comprises data slices from said second data layer that have been compressed.

20. A memory device having program instructions stored thereon that, upon execution by a computer system, cause the computer system to provide:

a first data storage layer comprising data blocks;

a second data storage layer comprising de-duplicated data evicted from the first data storage layer, wherein the second storage layer has a higher retrieval latency first data storage layer; and

a barrier storage area comprising a subset of the de-duplicated data that is stored for a settlement period after which the subset of the de-duplicated data is deleted, wherein the subset of the de-duplicated data is selected for transfer to a third storage layer operated by a cloud storage provider via a network, wherein the third data storage layer has a higher retrieval latency than the second data storage layer, wherein the settlement period is configurable dependent upon the cloud storage provider being used, wherein the settlement period represents a time after which the de-duplicated data stored by the cloud storage provider is made accessible, and wherein a given data item is stored in the first, second, or third data storage layer based upon frequency of use of the given data item.

Assignments (4)
MERGER Recorded Mar 26, 2019
From: STORSIMPLE, INC.
To: MICROSOFT CORPORATION
Reel/Frame 048700/0047 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 26, 2019
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 048700/0111 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME AND AN INVENTOR NAME PREVIOUSLY RECORDED ON REEL 029180 FRAME 0948. ASSIGNOR(S) HEREBY CONFIRMS THE STORSIMPLE IS NOW STORSIMPLE, INC. AND INVENTOR KULANGARA KURIAKOSE GEORGE IS NOW KURIAKOSE GEORGE KULANGARA. Recorded Nov 2, 2012
From: TESTARDI, RICHARD; COMETTO, MAURILIO; KULANGARA, KURIAKOSE GEORGE
To: STORSIMPLE, INC.
Reel/Frame 029237/0546 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2012
From: TESTARDI, RICHARD; COMETTO, MAURILIO; KULANGARA, KURIAKOSE GEORGE
To: STORSIMPLE
Reel/Frame 029180/0948 →
Continuity (2)
Continuation 12930502 · Jan 6, 2011
Related Publication 20130246711A1 · Sep 19, 2013