Systems and methods for data-aware storage tiering for deep learning
Systems and methods are configured to split an epoch associated with a training dataset into a plurality of mini-epochs. A machine learning model can be trained with a mini-epoch of the plurality of mini-epochs. The mini-epoch can be, during the training, iterated for a number of times during the training. One or more metrics reflective of at least one of: a training loss, training accuracy, or validation accuracy of the machine learning model associated with the mini-epoch can be received. Whether to terminate iterations of the mini-epoch early before a number of iterations of the mini-epoch reaches the number of times based on the one or more metrics can be determined. The number of iterations can be a non-zero number.
1 . A computer-implemented method comprising:
splitting, by a computing system, an epoch associated with a training dataset stored in a capacity tier of memory into a plurality of mini-epochs;
training, by the computing system, a machine learning model with a mini-epoch of the plurality of mini-epochs, wherein the mini-epoch is to be iterated during the training for a number of times reflected in a repeating factor, wherein a performance tier of the memory stores the mini-epoch;
receiving, by the computing system, one or more metrics reflective of at least one of: a training loss, training accuracy, or validation accuracy of the machine learning model associated with the mini-epoch;
determining, by the computing system, whether to terminate iterations of the mini-epoch early before a number of iterations of the mini-epoch reaches the number of times based on the one or more metrics, wherein the number of iterations is a non-zero number; and
prefetching a different mini-epoch from the capacity tier into a remaining portion of the performance tier unoccupied by the mini-epoch enabling execution of the number of iterations at the performance tier while the different mini-epoch is prefetched such that input-output stalls are avoided or reduced while maintaining machine learning model convergence by reducing input-output bandwidth demand in accordance with the repeating factor, and freeing the input-output bandwidth demand for nodes or applications sharing one of the capacity tier or the performance tier of the memory.
2 . The method of claim 1 , further comprising:
accessing the performance tier of memory that is coupled with at least one computing element designated to execute the training; and
accessing the capacity tier of memory that is coupled with the performance tier, wherein the performance tier has faster read throughput than the capacity tier.
3 . The method of claim 2 , wherein the mini-epoch has a size that is smaller than or equal to half of the total capacity of the performance tier.
4 . The method of claim 2 , wherein the at least one computing element includes at least one hardware accelerator that is coupled with the performance tier.
5 . The method of claim 2 , wherein:
the at least one computing element consumes data at an effective bandwidth rate (EB1),
the capacity tier provides data at a read throughput rate (B2), and
the number of times is greater than or equal to EB1 divided by B2 (EB1/B2).
6 . The method of claim 2 , further comprising:
determining that the machine learning model has been trained with the mini-epoch for the number of times; and
training the machine learning model with the different mini-epoch,
wherein during the training the machine learning model with the different mini-epoch, prefetching the next different mini-epoch into the performance tier unoccupied by the different mini-epoch.
7 . The method of claim 2 , further comprising:
calculating a score based on a combination of the one or more metrics,
wherein the determining whether to terminate the iterations of the mini-epoch early before the number of iterations of the mini-epoch reaches the number of times based on the one or more metrics comprises:
determining that the score does not improve during the training the machine learning model with the mini-epoch; and
terminating the training the machine learning model with the mini-epoch.
8 . The method of claim 7 , further comprising:
waiting until the prefetching the different mini-epoch completes;
after completion of the prefetching the different mini-epoch, training the machine learning model with the different mini-epoch.
9 . The method of claim 7 , further comprising:
determining that the score is dependent upon the number of times during the training of the machine learning model with the mini-epoch; and
adjusting the number of times based on the one or more metrics.
10 . The method of claim 1 , further comprising:
determining that more than a threshold level of bias is added during the training of the machine learning model with the mini-epoch; and
composing a different mini-epoch with random selection of training data in the different mini-epoch; or
increasing a size of the different mini-epoch.
11 . A system comprising:
at least one processor; and
a memory storing instructions that, when executed by the at least one processor, cause the system to perform a method comprising:
splitting, by a computing system, an epoch associated with a training dataset stored in a capacity tier of memory into a plurality of mini-epochs;
training a machine learning model with a mini-epoch of the plurality of mini-epochs, wherein the mini-epoch is to be iterated during the training for a number of times reflected in a repeating factor, wherein a performance tier of the memory stores the mini-epoch;
receiving one or more metrics reflective of at least one of: a training loss, training accuracy, or validation accuracy of the machine learning model associated with the mini-epoch;
determining whether to terminate iterations of the mini-epoch early before a number of iterations of the mini-epoch reaches the number of times based on the one or more metrics, wherein the number of iterations is a non-zero number; and
prefetching a different mini-epoch from the capacity tier into a remaining portion of the performance tier unoccupied by the mini-epoch enabling execution of the number of iterations at the performance tier while the different mini-epoch is prefetched such that input-output stalls are avoided or reduced while maintaining machine learning model convergence by reducing input-output bandwidth demand in accordance with the repeating factor, and freeing the input-output bandwidth demand for nodes or applications sharing one of the capacity tier or the performance tier of the memory.
12 . The system of claim 11 , wherein the instructions cause the system to perform the method further comprising:
accessing the performance tier of memory that is coupled with at least one computing element designated to execute the training; and
accessing the capacity tier of memory that is coupled with the performance tier, wherein the performance tier has faster read throughput than the capacity tier.
13 . The system of claim 12 , wherein the mini-epoch has a size that is smaller than or equal to half of the total capacity of the performance tier.
14 . The system of claim 12 , wherein the at least one computing element includes at least one hardware accelerator that is coupled with the performance tier.
15 . The system of claim 12 , wherein:
the at least one computing element consumes data at an effective bandwidth rate (EB1),
the capacity tier provides data at a read throughput rate (B2), and
the number of times is greater than or equal to EB1 divided by B2 (EB1/B2).
16 . A non-transitory computer-readable storage medium including instructions that, when executed by at least one processor of a computing system, cause the computing system to perform a method comprising:
training a machine learning model with a mini-epoch of the plurality of mini-epochs, wherein the mini-epoch is to be iterated during the training for a number of times reflected in a repeating factor, wherein individual mini-epochs of the plurality of mini-epochs being associated with a training dataset stored in a capacity tier of memory, wherein a performance tier of the memory stores the plurality of mini-epochs;
receiving one or more metrics reflective of at least one of: a training loss, training accuracy, or validation accuracy of the machine learning model associated with the mini-epoch;
determining whether to terminate iterations of the mini-epoch early before a number of iterations of the mini-epoch reaches the number of times based on the one or more metrics, wherein the number of iterations is a non-zero number; and
prefetching a different mini-epoch from the capacity tier into a remaining portion of the performance tier unoccupied by the mini-epoch enabling execution of the number of iterations at the performance tier while the different mini-epoch is prefetched such that input-output stalls are avoided or reduced while maintaining machine learning model convergence by reducing input-output bandwidth demand in accordance with the repeating factor, and freeing the input-output bandwidth demand for nodes or applications sharing one of the capacity tier or the performance tier of the memory.
17 . The non-transitory computer-readable storage medium of claim 16 , wherein the instructions cause the system to perform the method further comprising:
accessing the performance tier of memory that is coupled with at least one computing element designated to execute the training; and
accessing the capacity tier of memory that is coupled with the performance tier, wherein the performance tier has faster read throughput than the capacity tier.
18 . The non-transitory computer-readable storage medium of claim 17 , wherein the mini-epoch has a size that is smaller than or equal to half of the total capacity of the performance tier.
19 . The non-transitory computer-readable storage medium of claim 17 , wherein the at least one computing element includes at least one hardware accelerator that is coupled with the performance tier.
20 . The non-transitory computer-readable storage medium of claim 17 , wherein:
the at least one computing element consumes data at an effective bandwidth rate (EB1),
the capacity tier provides data at a read throughput rate (B2), and
the number of times is greater than or equal to EB1 divided by B2 (EB1/B2).