IP Library › Granted Patent US 12,645,931
Granted Patent B2
US 12,645,931 · App. 17/226,917 · Granted Jun 2, 2026

Systems and methods for data-aware storage tiering for deep learning

Inventors: Cong Xu (Milpitas, CA); Suparna Bhattacharya (Bangalore, IN); Paolo Faraboschi (Milpitas, CA)
Assignee: Hewlett Packard Enterprise Development LP
G06N3/08G06F16/217G06N3/049G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,931
App. No.
17/226,917
Granted
Jun 2, 2026
Kind
B2
Abstract

Systems and methods are configured to split an epoch associated with a training dataset into a plurality of mini-epochs. A machine learning model can be trained with a mini-epoch of the plurality of mini-epochs. The mini-epoch can be, during the training, iterated for a number of times during the training. One or more metrics reflective of at least one of: a training loss, training accuracy, or validation accuracy of the machine learning model associated with the mini-epoch can be received. Whether to terminate iterations of the mini-epoch early before a number of iterations of the mini-epoch reaches the number of times based on the one or more metrics can be determined. The number of iterations can be a non-zero number.

Claims (65)

1 . A computer-implemented method comprising:

splitting, by a computing system, an epoch associated with a training dataset stored in a capacity tier of memory into a plurality of mini-epochs;

training, by the computing system, a machine learning model with a mini-epoch of the plurality of mini-epochs, wherein the mini-epoch is to be iterated during the training for a number of times reflected in a repeating factor, wherein a performance tier of the memory stores the mini-epoch;

receiving, by the computing system, one or more metrics reflective of at least one of: a training loss, training accuracy, or validation accuracy of the machine learning model associated with the mini-epoch;

determining, by the computing system, whether to terminate iterations of the mini-epoch early before a number of iterations of the mini-epoch reaches the number of times based on the one or more metrics, wherein the number of iterations is a non-zero number; and

prefetching a different mini-epoch from the capacity tier into a remaining portion of the performance tier unoccupied by the mini-epoch enabling execution of the number of iterations at the performance tier while the different mini-epoch is prefetched such that input-output stalls are avoided or reduced while maintaining machine learning model convergence by reducing input-output bandwidth demand in accordance with the repeating factor, and freeing the input-output bandwidth demand for nodes or applications sharing one of the capacity tier or the performance tier of the memory.

2 . The method of claim 1 , further comprising:

accessing the performance tier of memory that is coupled with at least one computing element designated to execute the training; and

accessing the capacity tier of memory that is coupled with the performance tier, wherein the performance tier has faster read throughput than the capacity tier.

3 . The method of claim 2 , wherein the mini-epoch has a size that is smaller than or equal to half of the total capacity of the performance tier.

4 . The method of claim 2 , wherein the at least one computing element includes at least one hardware accelerator that is coupled with the performance tier.

5 . The method of claim 2 , wherein:

the at least one computing element consumes data at an effective bandwidth rate (EB1),

the capacity tier provides data at a read throughput rate (B2), and

the number of times is greater than or equal to EB1 divided by B2 (EB1/B2).

6 . The method of claim 2 , further comprising:

determining that the machine learning model has been trained with the mini-epoch for the number of times; and

training the machine learning model with the different mini-epoch,

wherein during the training the machine learning model with the different mini-epoch, prefetching the next different mini-epoch into the performance tier unoccupied by the different mini-epoch.

7 . The method of claim 2 , further comprising:

calculating a score based on a combination of the one or more metrics,

wherein the determining whether to terminate the iterations of the mini-epoch early before the number of iterations of the mini-epoch reaches the number of times based on the one or more metrics comprises:

determining that the score does not improve during the training the machine learning model with the mini-epoch; and

terminating the training the machine learning model with the mini-epoch.

8 . The method of claim 7 , further comprising:

waiting until the prefetching the different mini-epoch completes;

after completion of the prefetching the different mini-epoch, training the machine learning model with the different mini-epoch.

9 . The method of claim 7 , further comprising:

determining that the score is dependent upon the number of times during the training of the machine learning model with the mini-epoch; and

adjusting the number of times based on the one or more metrics.

10 . The method of claim 1 , further comprising:

determining that more than a threshold level of bias is added during the training of the machine learning model with the mini-epoch; and

composing a different mini-epoch with random selection of training data in the different mini-epoch; or

increasing a size of the different mini-epoch.

11 . A system comprising:

at least one processor; and

a memory storing instructions that, when executed by the at least one processor, cause the system to perform a method comprising:

splitting, by a computing system, an epoch associated with a training dataset stored in a capacity tier of memory into a plurality of mini-epochs;

training a machine learning model with a mini-epoch of the plurality of mini-epochs, wherein the mini-epoch is to be iterated during the training for a number of times reflected in a repeating factor, wherein a performance tier of the memory stores the mini-epoch;

receiving one or more metrics reflective of at least one of: a training loss, training accuracy, or validation accuracy of the machine learning model associated with the mini-epoch;

determining whether to terminate iterations of the mini-epoch early before a number of iterations of the mini-epoch reaches the number of times based on the one or more metrics, wherein the number of iterations is a non-zero number; and

prefetching a different mini-epoch from the capacity tier into a remaining portion of the performance tier unoccupied by the mini-epoch enabling execution of the number of iterations at the performance tier while the different mini-epoch is prefetched such that input-output stalls are avoided or reduced while maintaining machine learning model convergence by reducing input-output bandwidth demand in accordance with the repeating factor, and freeing the input-output bandwidth demand for nodes or applications sharing one of the capacity tier or the performance tier of the memory.

12 . The system of claim 11 , wherein the instructions cause the system to perform the method further comprising:

accessing the performance tier of memory that is coupled with at least one computing element designated to execute the training; and

accessing the capacity tier of memory that is coupled with the performance tier, wherein the performance tier has faster read throughput than the capacity tier.

13 . The system of claim 12 , wherein the mini-epoch has a size that is smaller than or equal to half of the total capacity of the performance tier.

14 . The system of claim 12 , wherein the at least one computing element includes at least one hardware accelerator that is coupled with the performance tier.

15 . The system of claim 12 , wherein:

the at least one computing element consumes data at an effective bandwidth rate (EB1),

the capacity tier provides data at a read throughput rate (B2), and

the number of times is greater than or equal to EB1 divided by B2 (EB1/B2).

16 . A non-transitory computer-readable storage medium including instructions that, when executed by at least one processor of a computing system, cause the computing system to perform a method comprising:

training a machine learning model with a mini-epoch of the plurality of mini-epochs, wherein the mini-epoch is to be iterated during the training for a number of times reflected in a repeating factor, wherein individual mini-epochs of the plurality of mini-epochs being associated with a training dataset stored in a capacity tier of memory, wherein a performance tier of the memory stores the plurality of mini-epochs;

receiving one or more metrics reflective of at least one of: a training loss, training accuracy, or validation accuracy of the machine learning model associated with the mini-epoch;

determining whether to terminate iterations of the mini-epoch early before a number of iterations of the mini-epoch reaches the number of times based on the one or more metrics, wherein the number of iterations is a non-zero number; and

prefetching a different mini-epoch from the capacity tier into a remaining portion of the performance tier unoccupied by the mini-epoch enabling execution of the number of iterations at the performance tier while the different mini-epoch is prefetched such that input-output stalls are avoided or reduced while maintaining machine learning model convergence by reducing input-output bandwidth demand in accordance with the repeating factor, and freeing the input-output bandwidth demand for nodes or applications sharing one of the capacity tier or the performance tier of the memory.

17 . The non-transitory computer-readable storage medium of claim 16 , wherein the instructions cause the system to perform the method further comprising:

accessing the performance tier of memory that is coupled with at least one computing element designated to execute the training; and

accessing the capacity tier of memory that is coupled with the performance tier, wherein the performance tier has faster read throughput than the capacity tier.

18 . The non-transitory computer-readable storage medium of claim 17 , wherein the mini-epoch has a size that is smaller than or equal to half of the total capacity of the performance tier.

19 . The non-transitory computer-readable storage medium of claim 17 , wherein the at least one computing element includes at least one hardware accelerator that is coupled with the performance tier.

20 . The non-transitory computer-readable storage medium of claim 17 , wherein:

the at least one computing element consumes data at an effective bandwidth rate (EB1),

the capacity tier provides data at a read throughput rate (B2), and

the number of times is greater than or equal to EB1 divided by B2 (EB1/B2).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 9, 2021
From: XU, CONG; BHATTACHARYA, SUPARNA; FARABOSCHI, PAOLO
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 055881/0634 →
Continuity (1)
Related Publication 20220327376A1 · Oct 13, 2022
References Cited (18)
US 10572800B2 · Wang et al. · 2020 [cited by applicant]
US 11704535B1 · Vemuri · 2023 [cited by examiner]
US 20170228639A1 · Hara et al. · 2017 [cited by applicant]
US 20190303765A1 · Gou et al. · 2019 [cited by applicant]
US 20200342307A1 · Gou et al. · 2020 [cited by applicant]
US 20220156276A1 · Dain · 2022 [cited by examiner]
Baptista et al., Raising the Abstraction Level of a Deep Learning Design on FPGAs, Nov. 2020. (Year: 2020). [cited by examiner]
Zhu et al., Entropy-Aware I/O Pipelining for Large-Scale Deep Learning on HPC Systems, 2018 IEEE International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication Systems, Sep. 2018. (Y… [cited by examiner]
Peng et al., Optimus: An Efficient Dynamic Resource Scheduler for Deep Learning Clusters, EuroSys '18: Proceedings of the Thirteenth EuroSys Conference, No. 3, pp. 1-14, Apr. 2018. (Year: 2018). [cited by examiner]
Nakandala, S. et al., “Cerebro: A Data System for Optimized Deep Learning Model Selection,” Proceedings of the VLDB Endowment, 2020, pp. 2159-2173, vol. 13, issue 11, https://dl.acm.org/doi/pdf/10.14778/3407790.3407816,… [cited by applicant]
Fox et al., “Learning Everywhere: Pervasive Machine Learning for Effective High-Performance Computation”, IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), Feb. 27, 2019, 9 pages. [cited by applicant]
Jae Duk Seo, “[Archived Post ] Andrew Ng: Deep Learning, Self-Taught Learning and Unsupervised Feature Learning”, available online at <https://medium.com/@SeoJaeDuk/archived-post-andrew-ng-deep-learning-self-taught-lear… [cited by applicant]
Katharopoulos et al., “Not All Samples Are Created Equal: Deep Learning with Importance Sampling”, Proceedings of the 35th International Conference on Machine Learning, 2018, 13 pages. [cited by applicant]
Kumar et al., “Quiver: An Informed Storage Cache for Deep Learning”, 18th USENIX Conference on File and Storage Technologies, Feb. 25-27, 2020, 15 pages. [cited by applicant]
Meng et al., “Convergence Analysis of Distributed Stochastic Gradient Descent with Shuffling”, 31st Conference on Neural Information Processing Systems (NIPS 2017), Sep. 29, 2017, 19 pages. [cited by applicant]
Mohan et al., “Analyzing and Mitigating Data Stalls in DNN Training”, Jan. 19, 2021, 22 pages. [cited by applicant]
Yang et al., “Accelerating Data Loading in Deep Neural Network Training”, IEEE International Conference on High Performance Computing, Data and Analytics (HiPC), Oct. 2, 2019, 11 pages. [cited by applicant]
Zhu et al., “Entropy-Aware I/O Pipelining for Large-Scale Deep Learning on HPC Systems”, IEEE International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication Systems, 2018, pp. 145-15… [cited by applicant]