IP Library Granted Patent US 10,747,778
Granted Patent B2
US 10,747,778 · App. 15/664,738 · Granted Aug 18, 2020

Replication of data using chunk identifiers

Inventors: Anirvan Duttagupta (San Jose, CA); Apurv Gupta (Bangalore, IN); Dinesh Pathak (Bangalore, IN)
Assignee: Cohesity, Inc.
G06F16/27G06F16/1752G06F16/1873G06F16/2246
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,747,778
App. No.
15/664,738
Granted
Aug 18, 2020
Kind
B2
Abstract

A data identifier for each data portion of a first group of different data portions of a first version of data is determined. The first version of the data is represented in a tree structure that references the determined data identifiers. A second version of the data is represented in a second tree structure using at least a portion of elements of the first tree structure of the first version. The second tree structure references one or more data identifiers of a portion of the second version of the data that is different from the first version of the data. The one or more data identifiers of the portion of the second version of the data that is different from the first version of the data are identified and sent. A response indicating which of the data portions corresponding to the sent one or more data identifiers are requested to be provided for replication is received.

Claims (45)

1. A method of replicating data, comprising:

determining, by a source system, one or more corresponding data identifiers for each data chunk of a first group of data chunks of a first version of file system data;

representing the first version of the file system data in a first tree structure that references the determined one or more corresponding data identifiers;

representing a second version of the file system data in a second tree structure using at least a portion of elements of the first tree structure of the first version, wherein the second tree structure references at least one or more data identifiers associated with one or more data chunks of the second version of the file system data that are not included in the first version of the file system data;

identifying and sending from the source system to a destination system the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data and corresponding chunk offsets for the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data, wherein a chunk offset indicates a portion of the data chunk to which a data identifier applies, wherein the data chunk has a corresponding file offset of the file system data, wherein in response to receiving the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data, the destination system:

identifies the one or more data chunks corresponding to the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not stored at the destination system based on the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data and the corresponding chunk offsets for the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data, and

sends to the source system a response indicating which of the one or more data chunks corresponding to the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data; and

receiving at the source system from the destination system the response indicating which of the one or more data chunks corresponding to the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data are requested to be provided for replication.

2. The method of claim 1 , further comprising sending to the destination system the one or more data chunks indicated in the response.

3. The method of claim 2 , further comprising:

receiving an indication of an interruption associated with sending the one or more data chunks indicated in the response;

determining a checkpoint associated with the sending the one or more data chunks indicated in the response; and

resuming, based on the checkpoint, the sending of the one or more data chunks indicated in the response.

4. The method of claim 2 , wherein one or more nodes are configured to send in parallel different portions of the one or more data chunks indicated in the response.

5. The method of claim 1 , wherein sending the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data includes sending a corresponding file offset associated with each of the one or more data identifiers.

6. The method of claim 1 , wherein identifying the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data includes traversing the second tree structure and the first tree structure.

7. The method of claim 6 , further comprising determining one or more leaf nodes that are shared between the second tree structure and the first tree structure.

8. The method of claim 7 , further comprising determining one or more leaf nodes that are not shared between the second tree structure and the first tree structure.

9. The method of claim 1 , wherein a first node of the second tree structure corresponds to a modification to a particular data chunk of the first group corresponding to a second node of the first tree structure.

10. The method of claim 9 , wherein the first node of the second tree structure is associated with a first data identifier of the modified particular data chunk and the second node of the first tree structure is associated with a second data identifier of the particular data chunk.

11. The method of claim 9 , wherein the first node of the second tree structure is associated with a first data identifier of the modified particular data chunk and a second data identifier is associated with the first version of the particular data chunk.

12. The method of claim 1 , wherein at least one node of the second tree structure references at least one node of the first tree structure.

13. The method of claim 1 , wherein the first tree structure comprises a root node, one or more intermediate nodes, and one or more leaf nodes.

14. The method of claim 13 , wherein at least one of the one or more leaf nodes stores the one or more data identifiers associated with one or more data chunks of the second version of the file system data that is not included in the first version of the file system data.

15. The method of claim 1 , wherein a remote site is configured to receive the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data and is configured to store the one or more data chunks that correspond to the at least one of the one or more data identifiers, wherein the one or more data chunks that correspond to the at least one of the one or more data identifiers is associated with the second version of the file system data that is different than the first version of the file system data.

16. A system for replicating data, comprising:

a processor; and

a memory coupled to the processor configured to provide the processor with instructions, which when executed cause the processor to:

determine one or more corresponding data identifiers for each data chunk of a first group of data chunks of a first version of file system data;

represent the first version of the file system data in a first tree structure that references the determined one or more corresponding data identifiers;

represent a second version of the file system data in a second tree structure using at least a portion of elements of the first tree structure of the first version, wherein the second tree structure references at least one or more data identifiers associated with one or more data chunks of the second version of the file system data that are not included in the first version of the file system data;

identify and send from the system to a destination system the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data and corresponding chunk offsets for the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data, wherein a chunk offset indicates a portion of the data chunk to which a data identifier applies, wherein the data chunk has a corresponding file offset of the file system data, wherein in response to receiving the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data, the destination system is configured to:

identify the one or more data chunks corresponding to the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not stored at the destination system based on the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data and the corresponding chunk offsets for the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data, and

sends to the system a response indicating which of the one or more data chunks corresponding to the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data; and

receive from the destination system the response indicating which of the one or more data chunks corresponding to the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data are requested to be provided for replication.

17. The system of claim 16 , wherein the processor is further configured to send to the destination system the one or more data chunks indicated in the response.

18. The system of claim 16 , wherein to identify the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data, the processor is further configured to traverse the second tree structure.

19. A computer program product for replicating data, the computer program product being embodied in a non-transitory computer readable storage medium and comprising computer instructions for:

determining, by a source system, one or more corresponding data identifiers for each data chunk of a first group of data chunks of a first version of file system data;

representing the first version of the file system data in a first tree structure that references the determined one or more corresponding data identifiers;

representing a second version of the file system data in a second tree structure using at least a portion of elements of the first tree structure of the first version, wherein the second tree structure references at least one or more data identifiers associated with one or more data chunks of the second version of the file system data that are not included in the first version of the file system data;

identifying and sending from the source system to a destination system the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data and corresponding chunk offsets for the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data, wherein a chunk offset indicates a portion of the data chunk to which a data identifier applies, wherein the data chunk has a corresponding file offset of the file system data, wherein in response to receiving the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data, the destination system:

identifies the one or more data chunks corresponding to the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not stored at the destination system based on the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data and the corresponding chunk offsets for the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data, and

sends to the source system a response indicating which of the one or more data chunks corresponding to the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data; and

receiving at the source system from the destination system the response indicating which of the one or more data chunks corresponding to the at least one or more data identifiers associated with the one or more data chunks of the second version of the file system data that are not included in the first version of the file system data are requested to be provided for replication.

Assignments (4)
TERMINATION AND RELEASE OF INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Dec 10, 2024
From: FIRST-CITIZENS BANK & TRUST COMPANY (AS SUCCESSOR TO SILICON VALLEY BANK)
To: COHESITY, INC.
Reel/Frame 069584/0498 →
SECURITY INTEREST Recorded Dec 9, 2024
From: VERITAS TECHNOLOGIES LLC; COHESITY, INC.
To: JPMORGAN CHASE BANK. N.A.
Reel/Frame 069890/0001 →
SECURITY INTEREST Recorded Sep 23, 2022
From: COHESITY, INC.
To: SILICON VALLEY BANK, AS ADMINISTRATIVE AGENT
Reel/Frame 061509/0818 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2017
From: DUTTAGUPTA, ANIRVAN; GUPTA, APURV; PATHAK, DINESH
To: COHESITY, INC.
Reel/Frame 043728/0750 →
Continuity (1)
Related Publication 20190034507A1 · Jan 31, 2019