IP Library › Granted Patent US 12,650,954
Granted Patent B2
US 12,650,954 · App. 18/893,043 · Granted Jun 9, 2026

System for deduplication in data storage environments

Inventors: Roopesh Chuggani (Jaipur, IN); Dinakaran Narayanan (Chennai, IN); Palak Sharma (Gurgaon, IN); Mathankumar Devarajan (San Jose, CA)
Assignee: NetApp, Inc.
G06F16/1748
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,650,954
App. No.
18/893,043
Granted
Jun 9, 2026
Kind
B2
Abstract

Systems, methods, and software are disclosed herein for identifying duplicate blocks of a storage system and deduplicating the storage system. In one example, a method of operating a computing device includes scanning first metadata of blocks of a container file of a virtual volume to generate a first log file including records of virtual volume block numbers (VVBNs) and fingerprints of the blocks; scanning second metadata of blocks of an active file system of the virtual volume to generate a second log file including records of VVBNs and file block numbers (FBNs) of the blocks; generating tuples based on merging the records of the first log file and the second records of the second log file according to the VVBNs; identifying duplications among the blocks based on the tuples; and deduplicating the blocks based on the duplications in the active file system identified based on the tuples.

Claims (71)

1 . A computing apparatus comprising:

one or more computer readable storage media;

one or more processors operatively coupled with the one or more computer readable storage media; and

program instructions stored on the one or more computer readable storage media that, when executed by the one or more processors, direct the computing apparatus to at least:

scan first metadata of blocks of a container file of a virtual volume to generate a first log file including records of virtual volume block numbers (VVBNs) and fingerprints of the blocks;

scan second metadata of blocks of an active file system of the virtual volume to generate a second log file including records of VVBNs and file block numbers (FBNs) of the blocks;

generate tuples based on merging the records of the first log file and the records of the second log file according to the VVBNs;

identify duplications among the blocks based on the tuples; and

deduplicate the blocks based on the duplications in the active file system identified based on the tuples.

2 . The computing apparatus of claim 1 , wherein scan the first metadata of the blocks of the container file, the program instructions direct the computing apparatus to:

for a given VVBN in a range of VVBNs of the container file:

if the given VVBN is identified in an existing database of fingerprints, obtain a fingerprint associated with the given VVBN from the existing database; and

if the given VVBN is not identified in the existing database of fingerprints, compute a fingerprint based on data of a block associated with the VVBN.

3 . The computing apparatus of claim 1 , wherein to scan the second metadata of the blocks of the active file system, the program instructions direct the computing apparatus to:

for a given block of the blocks of the active file system:

upon determining that data of the given block is compressed according to a format of a previous storage system:

decompress the data of the given block;

perform an inline deduplication, compression, and compaction of the data of the given block; and

generate a record of the VVBN and FBN of the given block.

4 . The computing apparatus of claim 1 , wherein a tuple comprises a fingerprint attribute, an index node value, an FBN attribute, and a VVBN attribute of a block.

5 . The computing apparatus of claim 4 , wherein to identify the duplications among the blocks based on the tuples, the program instructions direct the computing apparatus to:

aggregate the tuples based on the fingerprint attributes to form groups;

order the tuples in each of the groups based on a block type of the tuples;

identify dependencies among the tuples; and

identify the duplications based on the dependencies.

6 . The computing apparatus of claim 5 , wherein the block type comprises one of: snapshot block and active file system block.

7 . The computing apparatus of claim 1 , wherein to deduplicate the blocks based on the tuples, the program instructions direct the computing apparatus to, for a dependency of the dependencies, execute block sharing based on a type of the dependency.

8 . The computing apparatus of claim 7 , wherein a type of the dependency comprises one of: snapshot block to snapshot block and snapshot block to active file system block.

9 . A method of operating a computing device comprising:

scanning first metadata of blocks of a container file of a virtual volume to generate a first log file including records of virtual volume block numbers (VVBNs) and fingerprints of the blocks;

scanning second metadata of blocks of an active file system of the virtual volume to generate a second log file including records of VVBNs and file block numbers (FBNs) of the blocks;

generating tuples based on merging the records of the first log file and the records of the second log file according to the VVBNs;

identifying duplications among the blocks based on the tuples; and

deduplicating the blocks based on the duplications in the active file system identified based on the tuples.

10 . The method of claim 9 , wherein scanning the first metadata of the blocks of the container file comprises:

for a given VVBN in a range of VVBNs of the container file:

if the given VVBN is identified in an existing database of fingerprints, obtaining a fingerprint associated with the given VVBN from the existing database; and

if the given VVBN is not identified in the existing database of fingerprints, computing a fingerprint based on data of a block associated with the VVBN.

11 . The method of claim 9 , wherein scanning the second metadata of the blocks of the active file system comprises:

for a given block of the blocks of the active file system:

upon determining that data of the given block is compressed according to a format of a previous storage system:

decompressing the data of the given block;

performing an inline deduplication, compression, and compaction of the data of the given block; and

generating a record of the VVBN and FBN of the given block.

12 . The method of claim 9 , wherein a tuple comprises a fingerprint attribute, an index node value, an FBN attribute, and a VVBN attribute of a block.

13 . The method of claim 12 , wherein identifying the duplications among the blocks based on the tuples comprises:

aggregating the tuples based on the fingerprint attributes to form groups;

ordering the tuples in each of the groups based on a block type of the tuples;

identifying dependencies among the tuples; and

identifying the duplications based on the dependencies.

14 . The method of claim 13 , wherein the block type comprises one of: snapshot block and active file system block.

15 . The method of claim 9 , wherein deduplicating the blocks based on the tuples comprises, for a dependency of the dependencies, executing block sharing based on a type of the dependency.

16 . The method of claim 15 , wherein a type of the dependency comprises one of: snapshot block to snapshot block and snapshot block to active file system block.

17 . A method of operating a computing device, comprising:

scanning first metadata of blocks of a container file of a virtual volume to generate a first log file including records of virtual volume block numbers (VVBNs) and fingerprints of the blocks;

scanning second metadata of blocks of an active file system of the virtual volume to generate a second log file including records of VVBNs and file block numbers (FBNs) of the blocks;

generating tuples based on merging the records of the first log file and the records of the second log file according to the VVBNs;

aggregating the tuples based on fingerprint attributes of the tuples to form groups;

ordering the tuples in each of the groups based on a block type of the tuples; and

identifying dependencies among the tuples.

18 . The method of claim 17 , wherein scanning the first metadata of the blocks of the container file comprises:

for a given VVBN in a range of VVBNs of the container file:

if the given VVBN is identified in an existing database of fingerprints, obtaining a fingerprint associated with the given VVBN from the existing database; and

if the given VVBN is not identified in the existing database of fingerprints, computing a fingerprint based on data of a block associated with the VVBN.

19 . The method of claim 17 , wherein scanning the second metadata of the blocks of the active file system comprises:

for a given block of the blocks of the active file system:

upon determining that data of the given block is compressed according to a format of a previous storage system:

decompressing the data of the given block;

performing an inline deduplication, compression, and compaction of the data of the given block; and

generating a record of the VVBN and FBN of the given block.

20 . The method of claim 17 , wherein a tuple comprises a fingerprint attribute, an index node value, an FBN attribute, and a VVBN attribute of a block.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2024
From: CHUGGANI, ROOPESH; NARAYANAN, DINAKARAN; SHARMA, PALAK; DEVARAJAN, MATHANKUMAR
To: NETAPP, INC.
Reel/Frame 068663/0344 →
Continuity (1)
Related Publication 20260086982A1 · Mar 26, 2026
References Cited (22)
US 7747584B1 · Jernigan, IV · 2010 [cited by examiner]
US 8583893B2 · Pruthi · 2013 [cited by examiner]
US 8600949B2 · Periyagaram · 2013 [cited by examiner]
US 8793290B1 · Pruthi · 2014 [cited by examiner]
US 8898119B2 · Sharma · 2014 [cited by examiner]
US 9280457B2 · Goel · 2016 [cited by examiner]
US 9880762B1 · Armangau · 2018 [cited by examiner]
US 10521400B1 · Basov · 2019 [cited by examiner]
US 10817206B2 · Soukhman · 2020 [cited by examiner]
US 11288239B2 · Bafna · 2022 [cited by examiner]
US 11947497B2 · Qiu · 2024 [cited by examiner]
US 11994957B1 · Lewis · 2024 [cited by examiner]
US 12511262B1 · Shabi · 2025 [cited by examiner]
US 12517656B2 · Subramanian · 2026 [cited by examiner]
US 20120005450A1 · Bomma · 2012 [cited by examiner]
US 20200019310A1 · Faibish · 2020 [cited by examiner]
US 20220107916A1 · Curtis-Maury · 2022 [cited by examiner]
US 20220335027A1 · Subramanian Seshadri · 2022 [cited by examiner]
US 20220405254A1 · Shatsky · 2022 [cited by examiner]
US 20230244571A1 · Natanzon · 2023 [cited by examiner]
US 20240330127A1 · Lewis · 2024 [cited by examiner]
US 20250124004A1 · Jernigan, IV · 2025 [cited by examiner]