Dynamically modifying replication intervals based on accumulated data
Estimates of amounts of time to transfer data to be replicated based on sizes of expected data transfers resulting from accumulated data since a previous data transfer are monitored. A determination as to whether an estimated amount of time to transfer the data exceeds an interval for replicating the data is made. The interval is determined based on a recovery point objective (RPO) for replicating the data. In response to determining that the estimated amount of time to transfer the data exceeds the interval, a subsequent interval for replicating the data is determined that satisfies the RPO.
1 . A storage system, comprising:
a memory; and
a processing device, operatively coupled to the memory, the processing device configured to:
monitor estimates of amounts of time to transfer accumulated data to be replicated based on sizes of expected data transfers resulting from the accumulated data and a bandwidth of a network to transfer the accumulated data, wherein the estimates of the amounts of time are calculated through the processing device;
determine whether an estimated amount of time to transfer the data exceeds an interval for replicating the data, the interval determined based on a recovery point objective (RPO) for replicating the data;
in response to determining that the estimated amount of time to transfer the data exceeds the interval, determine a subsequent interval for replicating the data that satisfies the RPO; and
replicate the data via data transfers over the network according to the subsequent interval.
2 . The storage system of claim 1 , wherein the subsequent interval is determined based on an increased size of accumulated data for an upcoming data transfer.
3 . The storage system of claim 1 , wherein the subsequent interval is determined based on a decreased bandwidth of a network transferring the data.
4 . The storage system of claim 3 , wherein the storage system is a distributed storage system having a plurality of storage nodes, the plurality of storage nodes comprising managed flash storage devices.
5 . The storage system of claim 1 , wherein the subsequent interval comprises a reduced time interval for starting of a next data transfer.
6 . The storage system of claim 1 , wherein the processing device is further configured to:
determine when a next data transfer of replicated accumulated data occurs, based on determining a size of accumulated data of the storage system since a snapshot of the storage system for the previous data transfer.
7 . The storage system of claim 1 , wherein the processing device is further configured to:
execute a snapshot of the storage system to start replicating for a next data transfer, responsive to identifying a burst of accumulating data.
8 . The storage system of claim 1 , wherein the processing device is further configured to:
determine an overwrite pattern of accumulating data of the storage system, wherein the determining the subsequent interval is further based on the overwrite pattern.
9 . The storage system of claim 1 wherein the RPO is a maximum acceptable amount of data accumulation loss as measured by time.
10 . A method, comprising:
monitoring estimates of amounts of time to transfer accumulated data to be replicated within a storage system based on sizes of expected data transfers resulting from the accumulated data and a bandwidth of a network to transfer the accumulated data, wherein the estimates of the amounts of time are calculated through the processing device
determining, by a processing device of a storage controller, whether an estimated amount of time to transfer the data exceeds an interval for replicating the data, the interval determined based on a recovery point objective (RPO) for replicating the data;
in response to determining that the estimated amount of time to transfer the data exceeds the interval, determining a subsequent interval for replicating the data that satisfies the RPO; and
replicating the data via data transfers over the network according to the subsequent interval.
11 . The method of claim 10 , wherein the subsequent interval is determined based on an increased size of accumulated data for an upcoming data transfer.
12 . The method of claim 10 , wherein the subsequent interval is determined based on a decreased bandwidth of a network transferring the data.
13 . The method of claim 12 , wherein the storage system is a distributed storage system having a plurality of storage nodes, the plurality of storage nodes comprising managed flash storage devices.
14 . The method of claim 10 , wherein the subsequent interval comprises a reduced time interval for starting of a next data transfer.
15 . A non-transitory computer readable storage medium storing instructions which, when executed, cause a processing device to:
monitor estimates of amounts of time to transfer accumulated data to be replicated within a storage system based on sizes of expected data transfers resulting from the accumulated data and a bandwidth of a network to transfer the accumulated data, wherein the estimates of the amounts of time are calculated through the processing device;
determine whether an estimated amount of time to transfer the data exceeds an interval for replicating the data, the interval determined based on a recovery point objective (RPO) for replicating the data;
in response to determining that the estimated amount of time to transfer the data exceeds the interval, determine a subsequent interval for replicating the data that satisfies the RPO; and
replicating the data via data transfers over the network according to the subsequent interval.
16 . The non-transitory computer readable storage medium of claim 15 , wherein the subsequent interval is determined based on an increased size of accumulated data for an upcoming data transfer.
17 . The non-transitory computer readable storage medium of claim 15 , wherein the subsequent interval is determined based on a decreased bandwidth of a network transferring the data.
18 . The non-transitory computer readable storage medium of claim 17 , wherein the storage system is a distributed storage system having a plurality of storage nodes, the plurality of storage nodes comprising managed flash storage devices.
19 . The non-transitory computer readable storage medium of claim 15 , wherein the subsequent interval comprises a reduced time interval for starting of a next data transfer.
20 . The non-transitory computer readable storage medium of claim 15 , wherein the processing device is further configured to:
determine when a next data transfer of replicated accumulated data occurs, based on determining a size of accumulated data of a storage system since a snapshot of the storage system for the previous data transfer.