IP Library › Granted Patent US 12,554,590
Granted Patent B2
US 12,554,590 · App. 18/304,359 · Granted Feb 17, 2026

Deduplicating files across multiple storage tiers in a clustered file system network

Inventors: Abhinav Duggal (Fremont, CA); Philip Shilane (Newtown, PA); George Mathew (Belmont, CA); Chegu Vinod (San Jose, CA)
Assignee: Dell Products L.P.
G06F11/1453G06F11/1464G06F16/1748G06F2201/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,554,590
App. No.
18/304,359
Filed
Apr 21, 2023
Granted
Feb 17, 2026
Kind
B2
Art Unit
2165
USPC
707/654
Abstract

Embodiments are described for a system and method facilitating deduplication in a multi-tier storage system in which a file can have different portions written to different tiers. A process partition the data space of each tier to a number of similarity groups and distributes the similarity groups across file system services in a cluster. The distribution is done in such a way that for a given similarity group ID, the same file system service owns the similarity groups of every tier. This allows for efficient checks for deduplication as it can be done local to a node rather than requiring remote procedure calls.

Claims (19)

1 . A method for deduplicating data of data assets in a multi-tier network having a plurality of different storage devices in which a file can have different portions written to different tiers, comprising:

organizing the storage devices into a plurality of storage tiers based on respective operating characteristics;

mapping each storage tier to a respective Service Level Agreements (SLA) dictating storage requirements for each of the data assets to a backup program;

mapping each SLA to one or more tiers of the plurality of tiers based on the storage requirements of a respective SLA to the operating characteristics of each tier;

partitioning a data space of each tier into a plurality of similarity groups (simgroup), each simgroup having a unique assigned simgroup identifier (ID);

mapping data of the data assets to the similarity groups based on organization of the data among the tiers;

distributing the similarity groups across file system services of deduplication nodes in the network so that each file system service owns a same similarity group for all the tiers;

forming each similarity group by applying a mapping function comprising a hash of fingerprint data for first level data segments (L1) of a Merkle tree organizing the data;

assigning each file system service to one or more simgroups through a range of simgroup IDs;

checking, as a local operation performed at each node, data deduplication using a portion of a fingerprint index associated with a respective range of simgroup IDs by checking each similarity group against a mapping table routing a corresponding L1 data segment to a corresponding instance of a deduplication service of the each node; and

performing deduplication of the data of a similarity group during a backup on a respective deduplication node.

2 . The method of claim 1 wherein the storage requirements comprise backup and restore latencies, media availability, and cost, and further wherein the operating characteristics comprise throughput (input/output rate), latency, security, and availability.

3 . The method of claim 2 further comprising characterizing the tiers along a performance scale ranging from high performance to low performance for throughput versus cost of storage, and wherein the data assets comprise at least one of files, directories, Mtrees, and namespaces.

4 . The method of claim 3 wherein the tiers comprise one or more types of storage media selected from: hard disk drives (HDDs), solid state drives (SSDs), flash memory, and cloud storage, and wherein the performance, availability and cost characteristics are different for each type of storage media.

5 . The method of claim 4 wherein the backup software comprises part of a deduplication backup system performing backup and restore operations for nodes of the multi-tier network, and wherein the network comprises a Santorini clustered network.

6 . The method of claim 5 wherein the deduplication node executes deduplication and compression services that pack unique data segments, and writes data segments as an object in an object store.

7 . The method of claim 6 wherein an SLA attribute is used with the similarity group to send the file to the appropriate tier in the network to meet the SLA.

8 . The method of claim 1 wherein the data assets comprise at least one of files, directories, Mtrees, and namespaces, and further wherein the backup software comprises part of a deduplication backup system performing backup and restore operations for nodes of the multi-tier network.

9 . The method of claim 1 wherein each similarity group is calculated for an L1 data segment based on SHA1 fingerprints of L0 segments of the Merkle tree, and wherein the deduplication service deduplicates the L0 segments relative to other fingerprints within a same similarity group.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2023
From: DUGGAL, ABHINAV; SHILANE`, PHILIP; MATHEW, GEORGE; VINOD, CHEGU
To: DELL PRODUCTS L.P.
Reel/Frame 063394/0871 →
Continuity (1)
Related Publication 20240354201A1 · Oct 24, 2024
References Cited (16)
US 9047302B1 · Bandopadhyay · 2015 [cited by examiner]
US 10990518B1 · Wallace · 2021 [cited by examiner]
US 20120095968A1 · Gold · 2012 [cited by examiner]
US 20120117029A1 · Gold · 2012 [cited by examiner]
US 20150006835A1 · Oberhofer · 2015 [cited by examiner]
US 20160231928A1 · Lewis · 2016 [cited by examiner]
US 20160246528A1 · Colgrove · 2016 [cited by examiner]
US 20170344491A1 · Pandurangan · 2017 [cited by examiner]
US 20180129426A1 · Aron · 2018 [cited by examiner]
US 20190042146A1 · Wysoczanski · 2019 [cited by examiner]
US 20200125273A1 · Aron · 2020 [cited by examiner]
US 20200336387A1 · Suzuki · 2020 [cited by examiner]
US 20220092022A1 · Agarwal · 2022 [cited by examiner]
US 20230176743A1 · Dalmatov · 2023 [cited by examiner]
WO WO2021015636A1 · 2021 [cited by examiner]
NPL-Cao: Cao et. al., “TDDFS: A Tier-Aware Data Deduplication-Based File System”, ACM Transactions on Storage, vol. 15, No. 1 , Article 4., Feb. 2019. (Year: 2019). [cited by examiner]