Technique for improving oplog flushing
An improved flushing technique controls a draining speed of data from a temporary storage tier to a backend storage tier of a node so that draining logic does not overwhelm an input/output (I/O) workload that is serviced by the storage tiers. Illustratively, the temporary storage tier is a persistent write buffer embodied as an operations log (oplog) and the backend storage tier is persistent physical disk storage embodied as an extent store. The technique improves an oplog flushing algorithm by enabling control of the oplog draining speed (rate) to provide consistent performance when the I/O workload (e.g., a primary ingest I/O stream) is serviced by the extent store and/or oplog during one or more states (e.g., static inertia state, idle state and rebuild state) of the oplog.
1 . A non-transitory computer readable medium including program instructions for execution on a processor of a node, the program instructions configured to:
receive an input/output (I/O) workload at the node, the I/O workload including write accesses having data directed to a virtual disk (vdisk) of the node, the write access data cached in an operations log (oplog) of the node;
control draining of the cached data from the oplog to persistent storage of the node according to a state of the oplog;
in response to the oplog being in a static state, predict a peak storage usage of an amount of storage consumed by the oplog based on a rate of ingesting of the write access data and a rate of the draining of the cached data; and
regulate the draining of the cached data such that the amount of storage consumed by the oplog approaches a predetermined peak oplog storage usage based on a pre-configured amount of consumed storage at the persistent storage.
2 . The non-transitory computer readable medium of claim 1 , wherein the program instructions configured to predict the peak storage usage are further configured to measure (i) a draining rate from a last episode of the oplog, and (ii) a draining rate from episodes of the oplog other than the last episode, wherein the episodes are of a predetermined size that correspond to portions of the oplog.
3 . The non-transitory computer readable medium of claim 2 , wherein each draining rate corresponds to a count of nullifications of records of a respective episode, wherein each record includes the cached data from a corresponding I/O workload write access.
4 . The non-transitory computer readable medium of claim 1 , wherein the program instructions configured to predict the peak storage usage of the are further configured to measure a reference count of episodes to live ranges of the episodes based on a reference map, wherein the episodes are of a predetermined size that correspond to portions of the oplog.
5 . The non-transitory computer readable medium of claim 1 , wherein the program instructions configured to predict the peak storage usage of the oplog are further configured to measure a rate of fragment addition, wherein the fragment corresponds to a contiguous region of data.
6 . The non-transitory computer readable medium of claim 1 , wherein the program instructions are further configured to, in response to the oplog being in an idle state wherein the I/O workload lacks random write operations for an idle time period, regulate the draining of the cached data to increase a rate of flushing proportional to a length of the idle time period.
7 . The non-transitory computer readable medium of claim 6 , wherein the idle time period is a sliding window.
8 . The non-transitory computer readable medium of claim 6 , wherein the program instructions are further configured to, in response to the oplog being in the idle state, account for types of I/O workload operations including (i) sequential writes, (ii) random writes and (iii) reads, and wherein the draining of the cached data is paused during periods of I/O operations that impact the oplog including random writes.
9 . The non-transitory computer readable medium of claim 1 , wherein the program instructions are further configured to, in response to the oplog being in a rebuild state wherein the oplog and the persistent storage are being rebuilt, regulate the oplog draining such that completion of the rebuild of the oplog and the persistent storage occur in close temporal proximity.
10 . The non-transitory computer readable medium of claim 9 , wherein the program instructions are further configured to, in response to the oplog being in the rebuild state, predict a length of time for the oplog rebuild and configure a controller to regulate the draining of the oplog.
11 . The non-transitory computer readable medium of claim 10 , wherein the controller is a proportional-integral-derivative (PID) controller.
12 . The non-transitory computer readable medium of claim 1 , wherein the program instructions are further configured to reclaim storage from the oplog by garbage collecting the drained cached data.
13 . A method comprising:
receiving an input/output (I/O) workload at a compute node, the I/O workload including write accesses having data directed to a virtual disk (vdisk) of the node, the write access data cached in an operations log (oplog) of the node;
controlling draining of the cached data from the oplog to persistent storage of the node according to a state of the oplog;
in response to the oplog being in a static state, predicting a peak storage usage of an amount of storage consumed by the oplog based on a rate of ingesting of the write access data and a rate of the draining of the cached data; and
regulating the draining of the cached data such that the amount of storage consumed by the oplog approaches a predetermined peak oplog storage usage based on a pre-configured amount of consumed storage at the persistent storage.
14 . The method of claim 13 , wherein predicting the peak storage usage further comprises measuring (i) a draining rate from a last episode of the oplog, and (ii) a draining rate from episodes of the oplog other than the last episode, wherein the episodes are of a predetermined size that correspond to portions of the oplog.
15 . The method of claim 14 , wherein each draining rate corresponds to a count of nullifications of records of a respective episode, wherein each record includes the cached data from a corresponding I/O workload write access.
16 . The method of claim 13 , wherein predicting the peak storage usage further comprises measuring a reference count of episodes to live ranges of the episodes based on a reference map, wherein the episodes are of a predetermined size that correspond to portions of the oplog.
17 . The method of claim 13 , wherein predicting the peak storage usage of the oplog further comprises measuring a rate of fragment addition, wherein the fragment corresponds to a contiguous region of data.
18 . The method of claim 13 , further comprising, in response to the oplog being in an idle state wherein the I/O workload lacks random write operations for an idle time period, regulating the draining of the cached data to increase a rate of flushing proportional to a length of the idle time period.
19 . The method of claim 18 , wherein the idle time period is a sliding window.
20 . The method of claim 18 , further comprising, in response to the oplog being in the idle state, accounting for types of I/O workload operations including (i) sequential writes, (ii) random writes and (iii) reads, and wherein the draining of the cached data is paused during periods of I/O operations that impact the oplog including random writes.
21 . The method of claim 13 , further comprising, in response to the oplog being in a rebuild state wherein the oplog and the persistent storage are being rebuilt, regulating the oplog draining such that completion of the rebuild of the oplog and the persistent storage occur in close temporal proximity.
22 . The method of claim 21 , further comprising, in response to the oplog being in the rebuild state, predicting a length of time for the oplog rebuild and configuring a controller to regulate the draining of the oplog.
23 . The method of claim 22 , wherein the controller is a proportional-integral-derivative (PID) controller.
24 . The method of claim 13 , further comprising reclaiming storage from the oplog by garbage collecting the drained cached data.
25 . An apparatus comprising:
a node having a processor and persistent storage, wherein the processor is configured to execute program instructions configured to:
receive an input/output (I/O) workload at the node, the I/O workload including write accesses having data directed to a virtual disk (vdisk) of the node, the write access data cached in an operations log (oplog) of the node;
control draining of the cached data from the oplog to the persistent storage according to a state of the oplog;
in response to the oplog being in a static state, predict a peak storage usage of an amount of storage consumed by the oplog based on a rate of ingesting of the write access data and a rate of the draining of the cached data; and
regulate the draining of the cached data such that the amount of storage consumed by the oplog approaches a predetermined peak oplog storage usage based on a pre-configured amount of consumed storage at the persistent storage.
26 . The apparatus of claim 25 , wherein the program instructions configured to predict the peak storage usage are further configured to measure (i) a draining rate from a last episode of the oplog, and (ii) a draining rate from episodes of the oplog other than the last episode, wherein the episodes are of a predetermined size that correspond to portions of the oplog.
27 . The apparatus of claim 26 , wherein each draining rate corresponds to a count of nullifications of records of a respective episode, wherein each record includes the cached data from a corresponding I/O workload write access.
28 . The apparatus of claim 25 , wherein the program instructions configured to predict the peak storage usage are further configured to measure a reference count of episodes to live ranges of the episodes based on a reference map, wherein the episodes are of a predetermined size that correspond to portions of the oplog.
29 . The apparatus of claim 25 , wherein the program instructions configured to predict the peak storage usage of the oplog are further configured to measure a rate of fragment addition, wherein the fragment corresponds to a contiguous region of data.
30 . The apparatus of claim 25 , wherein the program instructions are further configured to, in response to the oplog being in an idle state wherein the I/O workload lacks random write operations for an idle time period, regulate the draining of the cached data to increase a rate of flushing proportional to a length of the idle time period.
31 . The apparatus of claim 30 , wherein the idle time period is a sliding window.
32 . The apparatus of claim 30 , wherein the program instructions are further configured to, in response to the oplog being in the idle state, account for types of I/O workload operations including (i) sequential writes, (ii) random writes and (iii) reads, and wherein the draining of the cached data is paused during periods of I/O operations that impact the oplog including random writes.
33 . The apparatus of claim 25 , wherein the program instructions are further configured to, in response to the oplog being in a rebuild state wherein the oplog and the persistent storage are being rebuilt, regulate the oplog draining such that completion of the rebuild of the oplog and the persistent storage occur in close temporal proximity.
34 . The apparatus of claim 33 , wherein the program instructions are further configured to, in response to the oplog being in the rebuild state, predict a length of time for the oplog rebuild and configure a controller to regulate the draining of the oplog.
35 . The apparatus of claim 34 , wherein the controller is a proportional-integral-derivative (PID) controller.
36 . The apparatus of claim 25 , wherein the program instructions are further configured to reclaim storage from the oplog by garbage collecting the drained cached data.