IP Library Granted Patent US 10,769,025
Granted Patent B2
US 10,769,025 · App. 16/427,791 · Granted Sep 8, 2020

Indexing a relationship structure of a filesystem

Inventors: Apurv Gupta (Bangalore, IN); Akshat Agarwal (Delhi, IN)
Assignee: Cohesity, Inc.
G06F11/1451G06F16/128G06F16/182G06F16/9027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,769,025
App. No.
16/427,791
Granted
Sep 8, 2020
Kind
B2
Abstract

One or more storage locations of file inodes in a data source to be backed up are identified. Filesystem metadata information is extracted from the one or more identified storage locations. At least one item of the extracted filesystem metadata information includes a reference to a parent inode. The extracted filesystem metadata information is stored in a data structure. The contents of the data structure are analyzed to index a relationship structure of file system contents of the data source.

Claims (37)

1. A method, comprising:

identifying one or more storage locations associated with a plurality of inodes in a data source to be backed up, wherein the one or more identified storage locations of the plurality of inodes include one or more data ranges in a disk file of the data source which correspond to the plurality of inodes;

extracting information associated with the plurality of inodes from the one or more identified storage locations of the data source to be backed up to a storage system, wherein at least one item of the extracted information associated with the plurality of inodes includes a reference to a parent inode, wherein extracting the information associated with the plurality of inodes from the one or more identified storage locations of the data source to be backed up to the storage system includes copying the extracted information associated with the plurality of inodes to a first storage tier of the storage system and storing the extracted information associated with the plurality of inodes in one or more data structures that are stored in the first storage tier of the storage system, wherein the one or more data structures include one or more key-value stores; and

analyzing contents of the one or more data structures to index a relationship structure of the plurality of inodes of the data source, wherein analyzing the contents of the one or more data structures includes:

scanning the one or more key-value stores; and

generating the index of the relationship structure of the plurality of inodes of the data source based on the scanning of the one or more key-value stores.

2. The method of claim 1 , wherein identifying the one or more storage locations associated with the plurality of inodes in the data source to be backed up includes issuing a plurality of concurrent read requests for the one or more identified storage locations.

3. The method of claim 1 , wherein the storage system is configured to access the data source to be backed up using a distributed file system protocol.

4. The method of claim 1 , wherein one of the one or more data structures corresponds to a plurality of files.

5. The method of claim 1 , wherein one of the one or more data structures corresponds to a plurality of directories.

6. The method of claim 1 , wherein the one or more data structures include a plurality of entries corresponding to the plurality of inodes, each entry associates an inode identifier with metadata information associated with an inode having the inode identifier.

7. The method of claim 6 , wherein the metadata information associated with the inode having the inode identifier includes the reference to a corresponding parent inode.

8. The method of claim 1 , wherein analyzing the contents of the one or more data structures includes

determining parent relationships of inodes associated with the one or more data structures.

9. The method of claim 8 , further comprising reconstructing a filesystem tree structure based on the determined parent relationships.

10. A computer program product, the computer program product being embodied in a non-transitory computer readable storage medium and comprising computer instructions for:

identifying one or more storage locations associated with a plurality of inodes in a data source to be backed up, wherein the one or more identified storage locations of the plurality of inodes include one or more data ranges in a disk file of the data source which correspond to the plurality of inodes;

extracting information associated with the plurality of inodes from the one or more identified storage locations of the data source to be backed up to a storage system, wherein at least one item of the extracted information associated with the plurality of inodes includes a reference to a parent inode, wherein extracting the information associated with the plurality of inodes from the one or more identified storage locations of the data source to be backed up to the storage system includes copying the extracted information associated with the plurality of inodes to a first storage tier of the storage system and storing the extracted information associated with the plurality of inodes in one or more data structures that are stored in the first storage tier of the storage system, wherein the one or more data structures include one or more key-value stores; and

analyzing contents of the one or more data structures to index a relationship structure of the plurality of inodes of the data source, wherein analyzing the contents of the one or more data structures includes:

scanning the one or more key-value stores; and

generating the index of the relationship structure of the plurality of inodes of the data source based on the scanning of the one or more key-value stores.

11. The computer program product of claim 10 , wherein to identify the one or more storage locations associated with the plurality of inodes in the data source to be backed up includes issuing a plurality of concurrent read requests for the one or more identified storage locations.

12. The computer program product of claim 10 , wherein the first storage tier of the storage system is comprised of one or more solid state drives.

13. The computer program product of claim 11 , wherein the storage system is configured to access the data source to be backed up using a distributed file system protocol.

14. The computer program product of claim 10 , wherein the one or more data structures include a plurality of entries corresponding to the plurality of inodes, each entry associated an inode identifier with metadata information associated with an inode having the inode identifier.

15. The computer program product of claim 14 , wherein the metadata information associated with the inode having the inode identifier includes the reference to a corresponding parent inode.

16. The computer program product of claim 10 , wherein analyzing the contents of the one or more data structures further comprises computer instructions for

determining parent relationships of inodes associated with the one or more data structures; and

reconstructing a filesystem tree structure based on the determined parent relationships.

17. A system, comprising:

a processor configured to:

identify one or more storage locations associated with a plurality of inodes in a data source to be backed up, wherein the one or more identified storage locations of the plurality of inodes include one or more data ranges in a disk file of the data source which correspond to the plurality of inodes;

extract information associated with the plurality of inodes from the one or more identified storage locations of the data source to be backed up to a storage system, wherein at least one item of the extracted information associated with the plurality of inodes includes a reference to a parent inode, wherein to extract the information associated with the plurality of inodes from the one or more identified storage locations of the data source to be backed up to the storage system, the processor is configured to copy the extracted information associated with the plurality of inodes to a first storage tier of the storage system and store the extracted information associated with the plurality of inodes in one or more data structures that are stored in the first storage tier of the storage system, wherein the one or more data structures include one or more key-value stores; and

analyze contents of the one or more data structures to index a relationship structure of the plurality of inodes of the data source, wherein to analyze the contents of the one or more data structures, the processor is configured to:

scan the one or more key-value stores; and

generate the index of the relationship structure of the plurality of inodes of the data source based on the scan of the one or more key-value stores; and

a memory coupled to the processor and configured to provide the processor with instructions.

Assignments (4)
TERMINATION AND RELEASE OF INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Dec 10, 2024
From: FIRST-CITIZENS BANK & TRUST COMPANY (AS SUCCESSOR TO SILICON VALLEY BANK)
To: COHESITY, INC.
Reel/Frame 069584/0498 →
SECURITY INTEREST Recorded Dec 9, 2024
From: VERITAS TECHNOLOGIES LLC; COHESITY, INC.
To: JPMORGAN CHASE BANK. N.A.
Reel/Frame 069890/0001 →
SECURITY INTEREST Recorded Sep 23, 2022
From: COHESITY, INC.
To: SILICON VALLEY BANK, AS ADMINISTRATIVE AGENT
Reel/Frame 061509/0818 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2019
From: GUPTA, APURV; AGARWAL, AKSHAT
To: COHESITY, INC.
Reel/Frame 049916/0599 →
Continuity (2)
Provisional Application 62793702 · Jan 17, 2019
Related Publication 20200233751A1 · Jul 23, 2020
Cited By (1)
US 12,566,743