IP Library Granted Patent US 11,153,380
Granted Patent B2
US 11,153,380 · App. 16/791,154 · Granted Oct 19, 2021

Continuous backup of data in a distributed data store

Inventors: Yan Valerie Leshinsky (Kirkland, WA); Lon Lundgren (Kirkland, WA); Raman Mittal (Seattle, WA); Stefano Stefani (Issaquah, WA)
Assignee: Amazon Technologies, Inc.
H04L67/1095H04L67/1076H04L67/1097
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,153,380
App. No.
16/791,154
Granted
Oct 19, 2021
Kind
B2
Abstract

A distributed data store may provide continuous backup for data stored in the distributed data store. Updates to data may be replicated amongst storage nodes according to a peer-to-peer replication scheme. A backup node may participate in the peer-to-peer replication scheme to identify additional updates to be applied to a backup version of the data in a separate data store. The backup node may obtain the updates according to the peer-to-peer replication scheme and update the backup version of the data. In some embodiments, configuration changes to the data in the distributed data store may be detected via the peer-to-peer replication scheme such that a backup node can adapt performance of backup operations in conformity with the configuration change.

Claims (49)

1. A system, comprising:

at least one processor; and

a memory, storing program instructions that when executed by the at least one processor, cause the at least one processor to:

detect, at a backup node that applies updates to a backup version of a data volume, a configuration change for the data volume, wherein the data volume is stored in a distributed data store across a plurality of storage nodes that provide the updates to the backup node according to a peer-to-peer replication scheme used to replicate updates to the data volume among the plurality of storage nodes and the backup node; and

adapt, by the backup node, performance of obtaining additional updates to the data volume according to the peer-to-peer replication scheme based, at least in part, on the detected configuration change.

2. The system of claim 1 ,

wherein the data configuration change is a change to the plurality of storage nodes that store the data volume; and

wherein to adapt the performance of obtaining additional updates to the data volume, the program instructions cause the at least one processor to identify a new storage node storing the data volume that participates in the peer-to-peer replication scheme.

3. The system of claim 2 , wherein the change to the plurality of storage nodes adds a storage node the plurality of storage nodes.

4. The system of claim 2 , wherein the change to the plurality of storage nodes replaces one of the plurality of storage nodes.

5. The system of claim 1 ,

wherein the data configuration change is a data size change of the data volume; and

wherein to adapt the performance of obtaining additional updates to the data volume, the program instructions cause the at least one processor to identify subsequent updates to be applied to the backup version of the according to a new task assignment that accounts for an additional backup node added for updating the backup version of the data volume.

6. The system of claim 1 ,

wherein the distributed data store is a log-structured data store;

wherein the data configuration change is a log truncation for the data volume; and

wherein to adapt the performance of obtaining additional updates to the data volume, the program instructions cause the at least one processor to obtain truncation information to include in the backup version of the data volume.

7. The system of claim 1 , wherein the data configuration change is detected based, at least in part, on a change to an identifier associated with a configuration of the data volume.

8. A method, comprising:

detecting, at a backup node that applies updates to a backup version of a data volume, a configuration change for the data volume, wherein the data volume is stored in a distributed data store across a plurality of storage nodes that provide the updates to the backup node according to a peer-to-peer replication scheme used to replicate updates to the data volume among the plurality of storage nodes and the backup node; and

adapting, by the backup node, performance of obtaining additional updates to the data volume according to the peer-to-peer replication scheme based, at least in part, on the detected configuration change.

9. The method of claim 8 ,

wherein the data configuration change is a change to the plurality of storage nodes that store the data volume; and

wherein adapting the performance of obtaining additional updates to the data volume comprises identifying a new storage node storing the data volume that participates in the peer-to-peer replication scheme.

10. The method of claim 9 , wherein the change to the plurality of storage nodes adds a storage node the plurality of storage nodes.

11. The method of claim 9 , wherein the change to the plurality of storage nodes replaces one of the plurality of storage nodes.

12. The method of claim 8 ,

wherein the data configuration change is a data size change of the data volume; and

wherein adapting the performance of obtaining additional updates to the data volume comprises identifying subsequent updates to be applied to the backup version of the according to a new task assignment that accounts for an additional backup node added for updating the backup version of the data volume.

13. The method of claim 8 ,

wherein the distributed data store is a log-structured data store;

wherein the data configuration change is a log truncation for the data volume; and

wherein adapting the performance of obtaining additional updates to the data volume comprises obtaining truncation information to include in the backup version of the data volume.

14. The method of claim 8 , wherein the data configuration change is detected based, at least in part, on a change to an identifier associated with a configuration of the data volume.

15. One or more non-transitory, computer-readable storage media, storing program instructions that when executed on or across one or more computing devices cause the one or more computing devices to implement:

detecting, at a backup node that applies updates to a backup version of a data volume, a configuration change for the data volume, wherein the data volume is stored in a distributed data store across a plurality of storage nodes that provide the updates to the backup node according to a peer-to-peer replication scheme used to replicate updates to the data volume among the plurality of storage nodes and the backup node; and

adapting, by the backup node, performance of obtaining additional updates to the data volume according to the peer-to-peer replication scheme based, at least in part, on the detected configuration change.

16. The one or more non-transitory, computer-readable storage media of claim 15 ,

wherein the data configuration change is a change to the plurality of storage nodes that store the data volume; and

wherein, in adapting the performance of obtaining additional updates to the data volume, the program instructions cause the one or more computing devices to implement identifying a new storage node storing the data volume that participates in the peer-to-peer replication scheme.

17. The one or more non-transitory, computer-readable storage media of claim 16 , wherein the change to the plurality of storage nodes adds a storage node the plurality of storage nodes.

18. The one or more non-transitory, computer-readable storage media of claim 16 , wherein the change to the plurality of storage nodes replaces one of the plurality of storage nodes.

19. The one or more non-transitory, computer-readable storage media of claim 15 ,

wherein the data configuration change is a data size change of the data volume; and

wherein, in adapting the performance of obtaining additional updates to the data volume, the program instructions cause the one or more computing devices to implement identifying subsequent updates to be applied to the backup version of the according to a new task assignment that accounts for an additional backup node added for updating the backup version of the data volume.

20. The one or more non-transitory, computer-readable storage media of claim 15 ,

wherein the distributed data store is a log-structured data store;

wherein the data configuration change is a log truncation for the data volume; and

wherein, in adapting the performance of obtaining additional updates to the data volume, the program instructions cause the one or more computing devices to implement obtaining truncation information to include in the backup version of the data volume.

Continuity (2)
Continuation 14977453 · Dec 21, 2015
Related Publication 20200186602A1 · Jun 11, 2020
Cited By (1)
US 12,197,288