IP Library Granted Patent US 10,380,072
Granted Patent B2
US 10,380,072 · App. 14/216,689 · Granted Aug 13, 2019

Managing deletions from a deduplication database

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,380,072
App. No.
14/216,689
Granted
Aug 13, 2019
Kind
B2
Abstract

An information management system can manage the removal of data block entries in a deduplicated data store using working copies of the data block entries residing in a local data store of a secondary storage computing device. The system can use the working copies to identify data blocks for removal. Once the deduplication database is updated with the changes to the working copies (e.g., using a transaction based update scheme), the system can query the deduplication database for the database entries identified for removal. Once identified, the system can remove the database entries identified for pruning and/or the corresponding deduplication data blocks from secondary storage.

Claims (32)

1. A method for removing information from a deduplication data store maintained in a secondary storage subsystem, the method comprising:

reviewing a plurality of working copies of data block entries residing in memory local to a secondary storage computing device to identify a first data block entry and a first data block corresponding to the first data block entry associated with a secondary storage operation, each of the plurality of working copies of data block entries corresponding to a data block entry stored in a first data store of a secondary storage subsystem that is distinct from the memory local to the secondary storage computing device,

the first data block being stored in a second data store of the secondary storage subsystem, the second data store storing a set of data blocks including the first data block and corresponding to a set of files formed from the set of data blocks and stored in deduplicated fashion,

the first data block entry being stored in the first data store of the secondary storage subsystem, the first data store storing a set of data block entries including the first data block entry, each entry in the set of data block entries corresponding to a respective data block in the set of data blocks and comprising at least a deduplication signature corresponding to the respective data block and a reference count corresponding to a number of instances of the respective data block included in the set of files;

modifying a working copy of the first data block entry residing in the memory local to the secondary storage computing device and corresponding to the first data block entry and the first data block;

updating the first data block entry stored in the first data store based on the modified working copy to indicate that the first data block should be removed from the second data store;

subsequent to said updating, querying the first data store to identify a group of one or more data blocks in the set of data blocks that should be removed from the second data store, the group including the first data block;

removing the group of one or more data blocks from the second data store; and

removing a group of one or more data block entries that correspond to the group of one or more data blocks from the first data store.

2. The method of claim 1 , wherein said updating comprises updating a value of the reference count of the first data block entry to generate a modified reference count value.

3. The method of claim 2 , wherein said querying comprises identifying the first data block for removal from the second data store based on the modified reference count value indicating that a number of references to the first data block is below a threshold value.

4. The method of claim 3 , wherein the threshold value is one and said updating of the value of the reference count of the first data block entry comprises decrementing the reference count of the first data block entry from one to zero.

5. The method of claim 1 , wherein said updating comprises setting a flag of the first data block entry.

6. The method of claim 1 , wherein said updating is initiated based on a detection of at least one of: expiration of a time threshold since a previous update to the first data store based on content of the memory local to the secondary storage computing device, and a size threshold of the memory local to the secondary storage computing device local data store being exceeded.

7. The method of claim 1 , wherein said updating comprises merging the modified working copy with the first data block entry contained in the first data store.

8. The method of claim 1 , wherein said removing the group of one or more data blocks from the second data store comprises, for at least one data block of the data blocks in the group, removing a copy of the data block from the second data store and additionally removing one or more pointers to the copy of the data block from the second data store.

9. A system for pruning a deduplication database, comprising:

a data block data store contained in one or more storage devices of a secondary storage subsystem, the data block data store storing a set of data blocks corresponding to a set of files formed from the set of data blocks, the set of files stored in deduplicated fashion;

a deduplication data store storing a set of data block entries, each entry of the set of data block entries corresponding to a respective data block of the set of data blocks and comprising at least a deduplication signature corresponding to the respective data block and a reference count corresponding to a number of instances of the respective data block included in the set of files; and

a secondary storage computing device residing in the secondary storage subsystem and comprising a local data store that is separate from the deduplication data store and resides in memory local to the secondary storage computing device, the local data store storing working copies of at least a subset of the set of data block entries, the secondary storage computing device further comprising computer hardware configured to:

review the working copies of the at least a subset of the set of data block entries to identify a first data block entry and a first data block corresponding to the first data block entry that are associated with a secondary storage operation, the first data block entry being stored in the deduplication data store and the first data block being stored in the data block data store;

modify a working copy of the first data block entry stored on the local data store that corresponds to the first data block entry and the first data block;

cause the deduplication data store to be updated based on the modified working copy to indicate that the first data block is to be removed from the data block data store;

query the deduplication data store to identify a group of one or more data blocks in the set of data blocks that are to be removed from the data block data store, the group including the first data block;

cause the group of one or more data blocks to be removed from data the data store; and

cause a group of one or more data block entries that correspond to the group of one or more data blocks to be removed from the deduplication data store.

10. The system of claim 9 , wherein to modify the working copy of the first data block entry, the secondary storage computing device is configured to modify a value of a reference count of the working copy of the first data block entry.

11. The system of claim 10 , wherein to identify the group of one or more data blocks for removal the secondary storage computing device is configured to identify the first data block for removal from the data block data store based on the modified value of the reference count indicating that a number of references to the first data block is below a threshold value.

12. The system of claim 11 , wherein the threshold value is one.

13. The system of claim 9 , wherein to modify the working copy of the first data block entry, the secondary storage computing device is configured to set a flag of the working copy of the first data block entry.

14. The system of claim 9 , wherein the secondary storage computing device is configured to cause the deduplication data store to be updated based on detection of at least one of: expiration of a time threshold since a previous update of the data block data store, and a size threshold of the local data store being exceeded.

15. The system of claim 9 , wherein an update of the secondary storage computing device includes merging content of the working copy of the first data block entry with the first data block entry contained in the deduplication data store.

Assignments (4)
SECURITY INTEREST Recorded Dec 13, 2021
From: COMMVAULT SYSTEMS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 058496/0836 →
RELEASE OF SECURITY INTEREST Recorded Jan 6, 2021
From: BANK OF AMERICA, N.A.
To: COMMVAULT SYSTEMS, INC.
Reel/Frame 054913/0905 →
SECURITY INTEREST Recorded Jul 2, 2014
From: COMMVAULT SYSTEMS, INC.
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 033266/0678 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2014
From: ATTARDE, DEEPAK RAGHUNATH; VIJAYAN, MANOJ KUMAR
To: COMMVAULT SYSTEMS, INC.
Reel/Frame 033056/0175 →
Cited By (2)
US 12,547,350 US 12,681,817