IP Library Granted Patent US 11,775,482
Granted Patent B2
US 11,775,482 · App. 16/854,153 · Granted Oct 3, 2023

File system metadata deduplication

Inventors: Anubhav Gupta (Sunnyvale, CA); Sachin Jain (Fremont, CA); Shreyas Talele (Santa Clara, CA); Zhihuan Qiu (Santa Clara, CA)
Assignee: Cohesity, Inc.
G06F16/174G06F16/14G06F16/184G06F16/9027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,775,482
App. No.
16/854,153
Granted
Oct 3, 2023
Kind
B2
Abstract

File metadata structures of a file system are analyzed. At least one metadata element that is duplicated among the analyzed file metadata structures is identified. The at least one identified metadata element is deduplicated including by modifying at least one of the file metadata structures to reference a same instance of the identified metadata element that is referenced by another one of the file metadata structures.

Claims (34)

1. A method, comprising:

analyzing file metadata structures of a file system, wherein each of the file metadata structures includes a plurality of metadata elements;

identifying a metadata element of the plurality of metadata elements that is duplicated among the analyzed file metadata structures, wherein instances of the identified metadata element reference a same data chunk, wherein the instances of the identified metadata element at least include a first leaf node of a first file metadata structure of the file metadata structures and a second leaf node of a second file metadata structure of the file metadata structures; and

deduplicating the instances of the identified metadata element at least in part by modifying the second file metadata structure to reference the first leaf node of the first file metadata structure of the file metadata structures instead of referencing the same data chunk,

wherein modifying the second file metadata structure to reference the first leaf node of the first file metadata structure includes modifying a parent metadata element of the second leaf node of the second file metadata structure to reference the first leaf node of the first file metadata structure instead of referencing the second leaf node of the second file metadata structure and deleting the second leaf node of the second file metadata structure, and

wherein the first leaf node of the first file metadata structure is an instance of the identified metadata element among the instances of the identified metadata element having a highest reference count.

2. The method of claim 1 , wherein analyzing file metadata structures of the file system comprises scanning a bottom level of the file metadata structures and a level above the bottom level of the file metadata structures.

3. The method of claim 1 , wherein at least one of the first file metadata structure and the second file metadata structure corresponds to a file generated by a storage cluster.

4. The method of claim 1 , wherein at least one of first file metadata structure and the second file metadata structure corresponds to a file backed up from a primary system to a storage cluster.

5. The method of claim 1 , further comprising identifying a plurality of file metadata structures from a set of stored file metadata structures that include the metadata element.

6. The method of claim 1 , wherein the metadata element is configured to store a value, wherein at least two of the analyzed file metadata structures include the metadata element that stores the value.

7. The method of claim 1 , wherein deduplicating instances of the identified metadata element includes identifying which instance of the identified metadata element has the highest reference count from a plurality of instances of the identified metadata element.

8. The method of claim 7 , wherein a reference count indicates a number of one or more other metadata elements that reference the metadata element.

9. The method of claim 1 , wherein deduplicating instances of the identified metadata element comprises deleting one or more instances of the identified metadata element not having the highest reference count.

10. The method of claim 9 , wherein the one or more instances of the identified metadata element not having the highest reference count are deleted from a solid state disk of a storage cluster.

11. The method of claim 1 , wherein the identified metadata element is deduplicated as a background process of a storage cluster.

12. The method of claim 11 , wherein the storage cluster is comprised of a plurality of storage nodes, wherein the metadata associated with a plurality of files is stored across the plurality of storage nodes.

13. The method of claim 1 , wherein deduplicating instances of the identified metadata element comprises determining a corresponding reference count associated with each instance of the identified metadata element.

14. The method of claim 13 , wherein the determined reference count is the same for each instance of the identified metadata element.

15. The method of claim 14 , wherein modifying the second file metadata structure includes selecting to reference the first leaf node of the first file metadata structure.

16. A computer program product, the computer program product being embodied in a non-transitory computer readable storage medium and comprising computer instructions for:

analyzing file metadata structures of a file system, wherein each of the file metadata structures includes a plurality of metadata elements;

identifying a metadata element of the plurality of metadata elements that is duplicated among the analyzed file metadata structures, wherein instances of the identified metadata element reference a same data chunk, wherein the instances of the identified metadata element at least include a first leaf node of a first file metadata structure of the file metadata structures and a second leaf node of a second file metadata structure of the file metadata structures; and

deduplicating the instances of the identified metadata element at least in part by modifying the second file metadata to reference the first leaf node of the first file metadata structure of the file metadata structures instead of referencing the same data chunk,

wherein modifying the second file metadata structure to reference the first leaf node of the first file metadata structure includes modifying a parent metadata element of the second leaf node of the second file metadata structure to reference the first leaf node of the first file metadata structure instead of referencing the second leaf node of the second file metadata structure and deleting the second leaf node of the second file metadata structure, and

wherein the first leaf node of the first file metadata structure is an instance of the identified metadata element among the instances of the identified metadata element having a highest reference count.

17. A system, comprising:

a processor configured to:

analyze file metadata structures of a file system, wherein each of the file metadata structures includes a plurality of metadata elements;

identify a metadata element of the plurality of metadata elements that is duplicated among the analyzed file metadata structures, wherein instances of the identified metadata element reference a same data chunk, wherein the instances of the identified metadata element at least include a first leaf node of a first file metadata structure of the file metadata structures and a second leaf node of a second file metadata structure of the file metadata structures; and

deduplicate the instances of the identified metadata element at least in part by modifying the second file metadata structure to reference the first leaf node of the first file metadata structure of the file metadata structures instead of referencing the same data chunk,

wherein modifying the second file metadata structure to reference the first leaf node of the first file metadata structure includes modifying a parent metadata element of the second leaf node of the second file metadata structure to reference the first leaf node of the first file metadata structure instead of referencing the second leaf node of the second file metadata structure and deleting the second leaf node of the second file metadata structure, and

wherein the first leaf node of the first file metadata structure is an instance of the identified metadata element among the instances of the identified metadata element having a highest reference count; and

a memory coupled to the processor and configured to provide the processor with instructions.

Assignments (4)
TERMINATION AND RELEASE OF INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Dec 10, 2024
From: FIRST-CITIZENS BANK & TRUST COMPANY (AS SUCCESSOR TO SILICON VALLEY BANK)
To: COHESITY, INC.
Reel/Frame 069584/0498 →
SECURITY INTEREST Recorded Dec 9, 2024
From: VERITAS TECHNOLOGIES LLC; COHESITY, INC.
To: JPMORGAN CHASE BANK. N.A.
Reel/Frame 069890/0001 →
SECURITY INTEREST Recorded Sep 23, 2022
From: COHESITY, INC.
To: SILICON VALLEY BANK, AS ADMINISTRATIVE AGENT
Reel/Frame 061509/0818 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 21, 2020
From: GUPTA, ANUBHAV; JAIN, SACHIN; TALELE, SHREYAS; QIU, ZHIHUAN
To: COHESITY, INC.
Reel/Frame 053266/0592 →
Continuity (2)
Provisional Application 62840614 · Apr 30, 2019
Related Publication 20200349115A1 · Nov 5, 2020
Cited By (1)
US 12,657,162