IP Library › Granted Patent US 10,642,794
Granted Patent B2
US 10,642,794 · App. 12/356,921 · Granted May 5, 2020

Computer storage deduplication

Inventors: Austin Clements (Cambridge, MA); Irfan Ahmad (Mountain View, CA); Jinyuan Li (Mountain View, CA); Murali Vilayannur (Sunnyvale, CA)
Assignee: VMware, Inc.
G06F16/1748
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,642,794
App. No.
12/356,921
Filed
Jan 21, 2009
Granted
May 5, 2020
Kind
B2
Examiner
ZHAO, YU
Art Unit
2169
USPC
707/2
Abstract

A data center comprising plural computer hosts and a storage system external to said hosts is disclosed. The storage system includes storage blocks for storing tangibly encoded data blocks. Each of said hosts includes a deduplicating file system for identifying and merging identical data blocks stored in respective storage blocks into one of said storage blocks so that a first file exclusively accessed by a first host of said hosts and a second file accessed exclusively by a second host of said hosts concurrently refer to the same one of said storage blocks.

Claims (47)

1. A data center comprising:

a plurality of host computers including a first host computer; and

a storage system external to and accessible by the plurality of host computers, wherein the storage system includes a plurality of storage blocks, a hash table, a write log and a merge log stored therein;

wherein each storage block in the plurality of storage blocks stores a data block and a reference count indicating a number of references in the storage system to the data block;

wherein the hash table contains hashes corresponding to used storage blocks, wherein a used storage block is a storage block with a reference count greater than zero;

wherein the write log contains write records, wherein each write record includes a reference to a storage block storing a data block written by the first host computer and a hash for the written data block; and

wherein the merge log is configured to store one or more merge requests;

wherein the first host computer is configured to:

retrieve one of the hashes from one of the write records in the write log;

determine a match between the retrieved hash and one of the hashes in the hash table for a used storage block other than the storage block storing the written data block corresponding to the retrieved hash;

determine that one of the plurality of host computers other than the first host computer has exclusive access to the storage block corresponding to the matching hash in the hash table, the other host having exclusive access by having a lock on a file containing the storage block; and

store a merge request in the merge log instead of performing a deduplication of the written data block and continue with deduplication operations on other files accessible to the first host computer, wherein the other host computer discovers the merge request stored in the merge log and based on the stored merge request performs the deduplication of the written data block by increasing the reference count for the storage block matching the hash in the hash table and freeing for reuse by the storage system the storage block containing the written data block.

2. The data center of claim 1 , wherein the first host is further configured to, after determining the match between the retrieved hash and one of the hashes in the hash table for the used storage block other than the storage block storing the written data block corresponding to the retrieved hash, compare data in the used storage block with data in the written data block bit-by-bit to confirm the match.

3. The data center of claim 1 , wherein the matching hash in the hash table is invalid due to a used storage block being overwritten after a hash record thereof is added to the hash table, and the deduplication is not performed.

4. The data center of claim 1 , wherein the hash table includes a plurality of hash records, each containing (1) a reference to a used storage block, and (2) a hash of the used storage block.

5. The data center of claim 1 ,

wherein the hash table is divided into a first portion accessible by the first host computer but not accessible by the other host computer, and a second portion accessible by the other host computer and not accessible by the first host computer; and

wherein the first host computer and the other host computer are able to access the respective portions of the hash table concurrently.

6. The data center of claim 1 , wherein the other computer performing the merge request further includes:

determining whether the retrieved hash matches one of the hashes in the hash table for a used block other than storage block storing the written data block corresponding to the retrieved hash; and

if the retrieved hash matches the one of the hashes in the hash table, performing the deduplication of the written data block.

7. The data center of claim 6 ,

wherein the storage system includes a pool of shared blocks; and

wherein the other computer performing the merge request further includes adding the storage block to the pool of the shared blocks.

8. A method for performing a deduplication operation in a storage system connected to a plurality of host computers including a first host computer, the storage system including a plurality of storage blocks a hash table, a write log, and a merge log stored therein,

wherein each storage block in the plurality of storage blocks stores a data block and a reference count indicating a number of references in the storage system to the data block,

wherein the hash table contains hashes corresponding to used storage blocks, wherein a used storage block is a storage block with a reference count greater than zero,

wherein the write log contains write records,

wherein each write record includes a reference to a storage block storage a data block written by the first host computer and a hash for the written data block, and

wherein the merge log is configured to store one or more merge requests;

the method comprising:

retrieving by the first host computer one of the hashes from one of the write records in the write log;

determining a match between the retrieved hash and one of the hashes in the hash table for a used storage block other than storage block storing the written data block corresponding to the retrieved hash;

determining that one of the plurality of host computers other than the first host computer has exclusive access to the storage block corresponding to the matching hash in the hash table, the other host having exclusive access by having a lock on a file containing the storage block; and

storing a merge request in the merge log instead of performing a deduplication of the written data block and continuing with deduplication operations on other files accessible to the first host computer, wherein the other host computer discovers the merge request stored in the merge log and based on the stored merge request performs the deduplication of the written data block by increasing the reference count for the storage block matching the hash in the hash table and freeing for reuse by the storage system the storage block containing the written data block.

9. The method of claim 8 , further comprising, after determining the match between the retrieved hash and one of the hashes in the hash table for a used storage block other than the storage block storing the written data block corresponding to the retrieved hash, comparing data in the used storage block with data in the written data block bit-by-bit to confirm the match.

10. The method of claim 8 , wherein the matching hash in the hash table is invalid due to a used storage block being overwritten after a hash record thereof is added to the hash table and the deduplication is not performed.

11. The method of claim 8 , wherein the hash table includes a plurality of hash records, each containing (1) a reference to a used storage block, and (2) a hash of the used storage block.

12. The method of claim 8 ,

wherein the hash table is divided into a first portion accessible by the first host computer but not accessible by the other host computer, and a second portion accessible by the other host computer and not accessible by the first host computer; and

wherein the first host computer and other host computer are able to access the respective portions of the hash table concurrently.

13. The method of claim 8 , wherein the other computer performing the merge request further includes:

determining whether the retrieved hash matches one of the hashes in the hash table for a used block other than storage block storing the written data block corresponding to the retrieved hash; and

if the retrieved hash matches the one of the hashes in the hash table, performing the deduplication of the written data block.

14. The method of claim 13 ,

wherein the storage system includes a pool of shared blocks; and

wherein the other computer performing the merge request further includes adding the storage block to the pool of the shared blocks.

Assignments (2)
CHANGE OF NAME Recorded Apr 15, 2024
From: VMWARE, INC.
To: VMWARE LLC
Reel/Frame 067102/0395 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 21, 2009
From: CLEMENTS, AUSTIN; AHMAD, IRFAN; LI, JINYUAN; VILAYANNUR, MURALI
To: VMWARE, INC.
Reel/Frame 022129/0620 →
Continuity (2)
Provisional Application 61096258 · Sep 11, 2008
Related Publication 20100077013A1 · Mar 25, 2010
Cited By (1)
US 12,360,967