IP Library › Granted Patent US 12,367,175
Granted Patent B2
US 12,367,175 · App. 17/951,241 · Granted Jul 22, 2025

Assessing the effectiveness of an archival job

Inventors: Shiv Kumar (Pune, IN); Kaushik Gupta (Pune, IN)
Assignee: Dell Products L.P.
G06F16/113G06F11/3409G06F16/125
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,175
App. No.
17/951,241
Granted
Jul 22, 2025
Kind
B2
Abstract

An archival job is assessed to calculate loss of data reduction efficiency due to block-level data deduplication. Archivable data, or individual storage objects or data structures therein, are moved to archival storage contingent upon satisfaction of a predetermined condition related to data reduction efficiency. Archivable data, or individual storage objects or data structures therein, that fail to satisfy the predetermined condition are maintained in primary storage. The loss of data reduction efficiency and the predetermined condition may be expressed as a percentage of maximum possible data reduction that would result in the absence of data deduplication.

Claims (43)

1. A method comprising:

in a storage system in which storage objects on primary storage archive to secondary storage only as single units and a first storage object contains archivable data, and a second storage object contains non-archivable data, wherein the archivable data and the non-archivable data share at least some duplicated data that has been consolidated into a single stored copy on the primary storage through deduplication:

identifying that the archivable data and the non-archivable data have been deduplicated to reference the single stored copy of the duplicated data;

determining, while data is static, whether archive of the first storage object is justified by:

calculating potential primary storage data reduction that would result from retaining the single stored copy of the duplicated data on primary storage and moving only a non-duplicated portion of the archivable data from the primary storage to archival storage; and

comparing the calculated potential primary storage data reduction that would result from retaining the single stored copy of the duplicated data on primary storage and moving only the non-duplicated portion of the archivable data from the primary storage to the archival storage with a predetermined condition;

retaining the archivable data on the primary storage without moving any of the archivable data to the archival storage in response to determining that the calculated potential primary storage data reduction fails to satisfy the predetermined condition due to the amount of duplication between the archivable data and the non-archivable data; and

copying the archivable data from the primary storage to the archival storage as the single unit and retaining the single stored copy of the duplicated data on primary storage in response to determining that the calculated potential primary storage data reduction satisfies the predetermined condition due to the amount of duplication between the archivable data and the non-archivable data.

2. The method of claim 1 further comprising identifying a set of data structures or the storage objects as archival candidates.

3. The method of claim 2 further comprising identifying inodes of the archival candidates and identifying blocks referenced by the inodes.

4. The method of claim 3 further comprising, for each referenced block, identifying an associated reference count of associations between the referenced block and the data structures or the storage objects on the primary storage.

5. The method of claim 4 further comprising incrementing a count of saved blocks in response to the reference count being equal to 1.

6. The method of claim 5 further comprising determining whether an entry for the referenced block has already been created in a hash table in response to the reference count being greater than 1.

7. The method of claim 6 further comprising creating an entry for the referenced block in the hash table in response to determining that no entry for the referenced block exists.

8. The method of claim 6 further comprising decrementing the reference count for the referenced block in the hash table in response to determining that an entry for the referenced block exists.

9. An apparatus comprising:

primary storage comprising non-volatile media on which is stored storage objects that archive to secondary storage only as single units;

a computer configured to manage access to data stored on the primary storage;

a first storage object that contains archivable data and a second storage object that contains non-archivable data stored on the primary storage, and wherein the archivable data and the non-archivable data share at least some duplicated data that has been consolidated into a single stored copy on the primary storage through deduplication; and

an archival data assessor comprising computer-executable instructions on a non-transitory computer-readable medium that determines, while data is static, whether archive of the first storage object is justified, the archival data assessor configured to:

identify that the archivable data and the non-archivable data have been deduplicated to reference the single stored copy of the duplicated data;

calculate potential primary storage data reduction that would result from retention of the single stored copy of the duplicated data on primary storage and movement of only a non-duplicated portion of the archivable data from the primary storage to archival storage;

compare the calculated potential primary storage data reduction that would result from retention of the single stored copy of the duplicated data on primary storage and movement of only the non-duplicated portion of the archivable data from the primary storage to the archival storage with a predetermined condition;

retain the archivable data on the primary storage without moving any of the archivable data to the archival storage in response to a determination that the calculated potential primary storage data reduction fails to satisfy the predetermined condition due to the amount of duplication of between the archivable data and the non-archivable data; and

copy the archivable data from the primary storage to the archival storage as the single unit and retain the single stored copy of the duplicated data on primary storage in response to determining that the calculated potential primary storage data reduction satisfies the predetermined condition due to the amount of duplication between the archivable data and the non-archivable data.

10. The apparatus of claim 9 further comprising the archival data assessor configured to identify a set of data structures or the storage objects as archival candidates.

11. The apparatus of claim 10 further comprising the archival data assessor configured to identify inodes of the archival candidates and identify blocks referenced by the inodes.

12. The apparatus of claim 11 further comprising the archival data assessor configured to, for each referenced block, identify an associated reference count of associations between the referenced block and the data structures or the storage objects on the primary storage.

13. The apparatus of claim 12 further comprising the archival data assessor configured to increment a count of saved blocks in response to the reference count being equal to 1.

14. The apparatus of claim 13 further comprising the archival data assessor configured to determine whether an entry for the referenced block has already been created in a hash table in response to the reference count being greater than 1.

15. The apparatus of claim 14 further comprising the archival data assessor configured to create an entry for the referenced block in the hash table in response to determining that no entry for the referenced block exists.

16. The apparatus of claim 14 further comprising the archival data assessor configured to decrement the reference count for the referenced block in the hash table in response to determining that an entry for the referenced block exists.

17. A non-transitory computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method comprising:

in a storage system in which storage objects on primary storage archive to secondary storage only as single units and a first storage object contains archivable data and a second storage object contains non-archivable data, wherein the archivable data and the non-archivable data share at least some duplicated data that has been consolidated into a single stored copy on the primary storage through deduplication:

identifying that archivable data and the non-archivable data have been deduplicated to reference the single stored copy of the duplicated data;

determining, while data is static, whether archive of the first storage object is justified by:

calculating potential primary storage data reduction that would result from retaining the single stored copy of the duplicated data on primary storage and moving only a non-duplicated portion of the archivable data from the primary storage to archival storage; and

comparing the calculated potential primary storage data reduction that would result from retaining the single stored copy of the duplicated data on primary storage and moving only the non-duplicated portion of the archivable data from the primary storage to the archival storage with a predetermined condition;

retaining the archivable data on the primary storage without moving any of the archivable data to the archival storage in response to determining that the calculated potential primary storage data reduction fails to satisfy the predetermined condition due to the amount of duplication between the archivable data and the non-archivable data; and

copying the archivable data from the primary storage to archival storage and retaining the single stored copy of the duplicated data on primary storage in response to determining that the calculated potential primary storage data reduction satisfies the predetermined condition due to the amount of duplication between the archivable data and the non-archivable data.

18. The non-transitory computer-readable storage medium of claim 17 in which the method further comprises identifying a set of data structures or the storage objects as archival candidates, identifying inodes of the archival candidates and identifying blocks referenced by the inodes, and, for each referenced block, identifying an associated reference count of associations between the referenced block and the data structures or the storage objects on the primary storage.

19. The non-transitory computer-readable storage medium of claim 18 in which the method further comprises incrementing a count of saved blocks in response to the reference count being equal to 1.

20. The non-transitory computer-readable storage medium of claim 19 in which the method further comprises determining whether an entry for the referenced block has already been created in a hash table in response to the reference count being greater than 1, creating an entry for the referenced block in the hash table in response to determining that no entry for the referenced block exists, and decrementing the reference count for the referenced block in the hash table in response to determining that an entry for the referenced block exists.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2022
From: KUMAR, SHIV; GUPTA, KAUSHIK
To: DELL PRODUCTS L.P.
Reel/Frame 061190/0927 →
Continuity (1)
Related Publication 20240104050A1 · Mar 28, 2024
References Cited (14)
US 7107298B2 · Prahlad · 2006 [cited by examiner]
US 9613064B1 · Chou · 2017 [cited by examiner]
US 10049116B1 · Bajpai · 2018 [cited by examiner]
US 10742735B2 · Kumar · 2020 [cited by examiner]
US 11561714B1 · Mertes · 2023 [cited by examiner]
US 20100333116A1 · Prahlad · 2010 [cited by examiner]
US 20120084523A1 · Littlefield · 2012 [cited by examiner]
US 20130275394A1 · Watanabe · 2013 [cited by examiner]
US 20130339407A1 · Sharpe · 2013 [cited by examiner]
US 20160036677A1 · McNutt · 2016 [cited by examiner]
US 20190065508A1 · Guturi · 2019 [cited by examiner]
US 20190213123A1 · Agarwal · 2019 [cited by examiner]
US 20230029616A1 · Pandit · 2023 [cited by examiner]
US 20230161672A1 · Raspudic · 2023 [cited by examiner]