IP Library Granted Patent US 11,693,741
Granted Patent B2
US 11,693,741 · App. 17/348,401 · Granted Jul 4, 2023

Large content file optimization

Inventors: Mohit Aron (Saratoga, CA); Zhihuan Qiu (San Jose, CA); Ganesha Shanmuganathan (San Jose, CA); Malini Mahalakshmi Venkatachari (Santa Clara, CA)
Assignee: Cohesity, Inc.
G06F11/1458G06F16/128G06F16/13G06F16/2246G06F2201/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,693,741
App. No.
17/348,401
Granted
Jul 4, 2023
Kind
B2
Abstract

A size associated with a content file is determined to be greater than a threshold size. Contents of the content file split across a plurality of component files are stored. Metadata, for the content file, is updated to reference a plurality of component file metadata structures for the component files. A node of the metadata is configured to track different sizes of portions of the content file stored in different component files of the plurality of component files. File metadata of the content file is split across the plurality of component file metadata structures and each component file metadata structure of the plurality of component file metadata structures specifies a corresponding structure organizing data components for a corresponding portion of the content file.

Claims (31)

1. A method, comprising:

performing a backup of a primary system that includes a plurality of portions of a content file that has a size that is greater than a threshold size;

storing the plurality of portions of the content file; and

generating a tree data structure that provides a view of the primary system, wherein generating the tree data structure includes generating a plurality of component file metadata structures for each of the plurality of portions of the content file, wherein a component file metadata structure of the plurality of component file metadata structures corresponds to one of the portions of the content file, wherein each of the plurality of component file metadata structures includes a corresponding root node, wherein each of the plurality of component file metadata structures includes metadata that enables data chunks associated with a corresponding portion of the content file to be located, wherein the tree data structure includes a plurality of leaf nodes, wherein a first leaf node of the plurality of leaf nodes stores a first vector that indicates a size of corresponding content file data that is associated with a corresponding component file metadata structure.

2. The method of claim 1 , wherein at least two of the plurality of portions of the content file have a same size.

3. The method of claim 1 , wherein at least two of the plurality of portions of the content file have a different size.

4. The method of claim 1 , further comprising determining that the size of the content file is greater than the threshold size.

5. The method of claim 1 , further comprising determining that the size of the content file is greater than the threshold size after performing the backup.

6. The method of claim 1 , wherein the component file metadata structure includes a second plurality of leaf nodes.

7. The method of claim 6 , wherein each of the second plurality of leaf nodes is associated with a corresponding data brick, wherein the corresponding data brick is an identifier for one or more data chunks.

8. The method of claim 7 , wherein a last data brick of the plurality of leaf nodes has a particular capacity, wherein the last data brick is brick aligned in the event the last data brick is associated with one or more data chunks having the particular capacity.

9. The method of claim 7 , wherein a last data brick of the plurality of leaf nodes has a particular capacity, wherein in the event the last data brick of the plurality of leaf nodes is not brick aligned, an unused portion of the last data brick is reserved for the content file.

10. The method of claim 1 , wherein the first leaf node of the plurality of leaf nodes stores information that indicates which component file metadata structure of the plurality of component file metadata structures is associated with which portion of the content file.

11. The method of claim 1 , wherein a plurality of sequential component file metadata structures associated with the content file have a same corresponding size.

12. The method of claim 11 , wherein the first vector utilizes run length encoding for the plurality of sequential component file metadata structures associated with the content file that have the same corresponding size.

13. The method of claim 12 , wherein the first leaf node of the plurality of leaf nodes stores a second vector that indicates a number of the sequential component file metadata structures that have the same corresponding size.

14. A computer program product, the computer program product being embodied in a non-transitory computer readable storage medium and comprising computer instructions for:

performing a backup of a primary system that includes a plurality of portions of a content file that has a size that is greater than a threshold size;

storing plurality of portions of content file; and

generating a tree data structure that provides a view of the primary system, wherein generating the tree data structure includes generating a plurality of component file metadata structures for each of the plurality of portions of the content file, wherein a component file metadata structure of the plurality of component file metadata structures corresponds to one of the portions of the content file, wherein each of the plurality of component file metadata structures includes a corresponding root node, wherein each of the plurality of component file metadata structures includes metadata that enables data chunks associated with a corresponding portion of the content file to be located, wherein the tree data structure includes a plurality of leaf nodes, wherein a first leaf node of the plurality of leaf nodes stores a first vector that indicates a size of corresponding content file data that is associated with a corresponding component file metadata structure.

15. The computer program product of claim 14 , further comprising determining that the size of the content file is greater than the threshold size.

16. The computer program product of claim 14 , further comprising determining that the size of the content file is greater than the threshold size after performing the backup.

17. The computer program product of claim 14 , wherein a plurality of sequential component file metadata structures associated with the content file have a same corresponding size.

18. The computer program product of claim 17 , wherein the first vector utilizes run length encoding for the plurality of sequential component file metadata structures associated with the content file that have the same corresponding size.

19. The computer program product of claim 18 , wherein the first leaf node of the plurality of leaf nodes stores a second vector that indicates a number of the sequential component file metadata structures that have the same corresponding size.

20. A system, comprising:

a processor configured to:

perform a backup of a primary system that includes a plurality of portions of a content file that has a size that is greater than a threshold size;

store the plurality of portions of the content file; and

generate a tree data structure that provides a view of the primary system, wherein to generate the tree data structure, the processor is configured to generate a plurality of component file metadata structures for each of the plurality of portions of the content file, wherein a component file metadata structure of the plurality of component file metadata structures corresponds to one of the portions of the content file, wherein each of the plurality of component file metadata structures includes a corresponding root node, wherein each of the plurality of component file metadata structures includes metadata that enables data chunks associated with a corresponding portion of the content file to be located, wherein the tree data structure includes a plurality of leaf nodes, wherein a first leaf node of the plurality of leaf nodes stores a first vector that indicates a size of corresponding content file data that is associated with a corresponding component file metadata structure; and

a memory coupled to the processor and configured to provide the processor with instructions.

Assignments (4)
TERMINATION AND RELEASE OF INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Dec 10, 2024
From: FIRST-CITIZENS BANK & TRUST COMPANY (AS SUCCESSOR TO SILICON VALLEY BANK)
To: COHESITY, INC.
Reel/Frame 069584/0498 →
SECURITY INTEREST Recorded Dec 9, 2024
From: VERITAS TECHNOLOGIES LLC; COHESITY, INC.
To: JPMORGAN CHASE BANK. N.A.
Reel/Frame 069890/0001 →
SECURITY INTEREST Recorded Sep 23, 2022
From: COHESITY, INC.
To: SILICON VALLEY BANK, AS ADMINISTRATIVE AGENT
Reel/Frame 061509/0818 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2021
From: ARON, MOHIT; QIU, ZHIHUAN; SHANMUGANATHAN, GANESHA; VENKATACHARI, MALINI MAHALAKSHMI
To: COHESITY, INC.
Reel/Frame 057316/0629 →
Continuity (3)
Continuation 16688653 · Nov 19, 2019
Continuation In Part 16024107 · Jun 29, 2018
Related Publication 20210382792A1 · Dec 9, 2021