IP Library › Granted Patent US 12,579,434
Granted Patent B2
US 12,579,434 · App. 17/865,029 · Granted Mar 17, 2026

Training a neural network using an accelerated gradient with shuffling

Inventors: Lam Minh Nguyen (Ossining, NY); Huyen Trang Tran (Ithaca, NY)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,434
App. No.
17/865,029
Granted
Mar 17, 2026
Kind
B2
Abstract

An index sequence specifying an index of training data corresponding to a component of a cost function is generated. A first model parameter in the set of model parameters is set to an initial value. Using the index sequence, a neural network model comprising a set of weights is trained. As part of the training, using the index sequence, a learning rate, and a set of gradients, a subset of the set of model parameters is updated. As part of the training, a momentum term is set. As part of the training, using the momentum term as the first model parameter, the updating and the setting are repeated until reaching a training completion condition. The trained neural network model is used to predict an outcome by analyzing live data.

Claims (66)

1 . A computer-implemented method comprising:

constructing a neural network training system with an improved performance in a training of a neural network model, the constructing comprising:

generating an index sequence of a set of training data, the index sequence specifying an index of training data corresponding to a component of a cost function;

setting a first model parameter in a set of model parameters to an initial value;

training, using the index sequence, the neural network model comprising a set of weights, the training comprising:

executing code for updating, using the index sequence, a learning rate, and a set of gradients, a subset of the set of model parameters, the subset comprising the set of model parameters excluding the first model parameter;

executing code for setting, using a difference between an updated value of a second model parameter in the set of model parameters and a previous value of the second model parameter, a momentum term;

executing code for repeating, using the momentum term as the first model parameter, the updating and the setting, the repeating performed until reaching a training completion condition, to produce a trained neural network model with a faster convergence rate of the cost function in the training relative to a second convergence rate of the cost function in a second training of the neural network model using stochastic gradient descent; and

using, to predict an outcome by analyzing live data, the trained neural network model.

2 . The computer-implemented method of claim 1 , wherein the index sequence is predetermined.

3 . The computer-implemented method of claim 1 , wherein the index sequence is generated, using a pseudo-random number generator, prior to the training.

4 . The computer-implemented method of claim 1 , wherein the index sequence is generated, using a pseudo-random number generator, prior to each iteration of the training.

5 . The computer-implemented method of claim 1 , wherein a model parameter in the set of model parameters comprises a value of a weight in the set of weights.

6 . The computer-implemented method of claim 1 , wherein the cost function satisfies a property in which a set of changes in the cost function between current values of the set of model parameters and an optimal set of model parameters is bounded within a threshold of a convex evaluation term and a squared distance between current values of the set of model parameters and an optimal set of model parameters, wherein the optimal set of model parameters comprises a set of predetermined values.

7 . The computer-implemented method of claim 1 , wherein the training completion condition comprises a change in an average of values of the set of model parameters that is less than a threshold amount.

8 . The computer-implemented method of claim 1 , wherein the momentum term comprises a sum of the updated value of the second model parameter and a difference between the updated value of the second model parameter and a value of the second model parameter for a previous epoch, the difference multiplied by a predetermined value.

9 . A computer program product for training a neural network, the computer program product comprising:

one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the stored program instructions comprising:

program instructions to construct a neural network training system with an improved performance in a training of a neural network model, the program instructions to construct comprising:

program instructions to generate an index sequence of a set of training data, the index sequence specifying an index of training data corresponding to a component of a cost function;

program instructions to set a first model parameter in a set of model parameters to an initial value;

program instructions to train, using the index sequence, the neural network model comprising a set of weights, the training comprising:

program instructions to update, using the index sequence, a learning rate, and a set of gradients, a subset of the set of model parameters, the subset comprising the set of model parameters excluding the first model parameter;

program instructions to set, using a difference between an updated value of a second model parameter in the set of model parameters and a previous value of the second model parameter, a momentum term;

program instructions to repeat, using the momentum term as the first model parameter, the updating and the setting, the repeating performed until reaching a training completion condition, to produce a trained neural network model with a faster convergence rate of the cost function in the training relative to a second convergence rate of the cost function in a second training of the neural network model using stochastic gradient descent; and

program instructions to use, to predict an outcome by analyzing live data, the trained neural network model.

10 . The computer program product of claim 9 , wherein the index sequence is predetermined.

11 . The computer program product of claim 9 , wherein the index sequence is generated, using a pseudo-random number generator, prior to the training.

12 . The computer program product of claim 9 , wherein the index sequence is generated, using a pseudo-random number generator, prior to each iteration of the training.

13 . The computer program product of claim 9 , wherein a model parameter in the set of model parameters comprises a value of a weight in the set of weights.

14 . The computer program product of claim 9 , wherein the cost function satisfies a property in which a set of changes in the cost function between current values of the set of model parameters and an optimal set of model parameters is bounded within a threshold of a convex evaluation term and a squared distance between current values of the set of model parameters and an optimal set of model parameters, wherein the optimal set of model parameters comprises a set of predetermined values.

15 . The computer program product of claim 9 , wherein the training completion condition comprises a change in an average of values of the set of model parameters that is less than a threshold amount.

16 . The computer program product of claim 9 , wherein the momentum term comprises a sum of the updated value of the second model parameter and a difference between the updated value of the second model parameter and a value of the second model parameter for a previous epoch, the difference multiplied by a predetermined value.

17 . The computer program product of claim 9 , wherein the stored program instructions are stored in the at least one of the one or more storage media of a local data processing system, and wherein the stored program instructions are transferred over a network from a remote data processing system.

18 . The computer program product of claim 9 , wherein the stored program instructions are stored in the at least one of the one or more storage media of a server data processing system, and wherein the stored program instructions are downloaded over a network to a remote data processing system for use in a computer readable storage device associated with the remote data processing system.

19 . The computer program product of claim 9 , wherein the computer program product is provided as a service in a cloud environment.

20 . A computer system comprising one or more processors, one or more computer-readable memories, and one or more computer-readable storage media, and program instructions stored on at least one of the one or more storage media for execution by at least one of the one or more processors via at least one of the one or more memories, the stored program instructions comprising:

program instructions to construct a neural network training system with an improved performance in a training of a neural network model, the program instructions to construct comprising:

program instructions to generate an index sequence of a set of training data, the index sequence specifying an index of training data corresponding to a component of a cost function;

program instructions to set a first model parameter in a set of model parameters to an initial value;

program instructions to train, using the index sequence, the neural network model comprising a set of weights, the training comprising:

program instructions to update, using the index sequence, a learning rate, and a set of gradients, a subset of the set of model parameters, the subset comprising the set of model parameters excluding the first model parameter;

program instructions to set, using a difference between an updated value of a second model parameter in the set of model parameters and a previous value of the second model parameter, a momentum term;

program instructions to repeat, using the momentum term as the first model parameter, the updating and the setting, the repeating performed until reaching a training completion condition, to produce a trained neural network model with a faster convergence rate of the cost function in the training relative to a second convergence rate of the cost function in a second training of the neural network model using stochastic gradient descent; and

program instructions to use, to predict an outcome by analyzing live data, the trained neural network model.

21 . The computer system of claim 20 , wherein the index sequence is predetermined.

22 . The computer system of claim 20 , wherein the index sequence is generated, using a pseudo-random number generator, prior to the training.

23 . The computer system of claim 20 , wherein the index sequence is generated, using a pseudo-random number generator, prior to each iteration of the training.

24 . A data processing environment comprising one or more processors, one or more computer-readable memories, and one or more computer-readable storage media, and program instructions stored on at least one of the one or more storage media for execution by at least one of the one or more processors via at least one of the one or more memories, the stored program instructions comprising:

program instructions to construct a neural network training system with an improved performance in a training of a neural network model, the program instructions to construct comprising:

program instructions to generate an index sequence of a set of training data, the index sequence specifying an index of training data corresponding to a component of a cost function;

program instructions to set a first model parameter in a set of model parameters to an initial value;

program instructions to train, using the index sequence, the neural network model comprising a set of weights, the training comprising:

program instructions to update, using the index sequence, a learning rate, and a set of gradients, a subset of the set of model parameters, the subset comprising the set of model parameters excluding the first model parameter;

program instructions to set, using a difference between an updated value of a second model parameter in the set of model parameters and a previous value of the second model parameter, a momentum term;

program instructions to repeat, using the momentum term as the first model parameter, the updating and the setting, the repeating performed until reaching a training completion condition, to produce a trained neural network model with a faster convergence rate of the cost function in the training relative to a second convergence rate of the cost function in a second training of the neural network model using stochastic gradient descent; and

program instructions to use, to predict an outcome by analyzing live data, the trained neural network model.

25 . A neural network model training system comprising one or more processors, one or more computer-readable memories, and one or more computer-readable storage media, and program instructions stored on at least one of the one or more storage media for execution by at least one of the one or more processors via at least one of the one or more memories, the stored program instructions comprising:

program instructions to construct a neural network training system with an improved performance in a training of a neural network model, the program instructions to construct comprising:

program instructions to generate an index sequence of a set of training data, the index sequence specifying an index of training data corresponding to a component of a cost function;

program instructions to set a first model parameter in a set of model parameters to an initial value;

program instructions to train, using the index sequence, the neural network model comprising a set of weights, the training comprising:

program instructions to update, using the index sequence, a learning rate, and a set of gradients, a subset of the set of model parameters, the subset comprising the set of model parameters excluding the first model parameter;

program instructions to set, using a difference between an updated value of a second model parameter in the set of model parameters and a previous value of the second model parameter, a momentum term;

program instructions to repeat, using the momentum term as the first model parameter, the updating and the setting, the repeating performed until reaching a training completion condition, to produce a trained neural network model with a faster convergence rate of the cost function in the training relative to a second convergence rate of the cost function in a second training of the neural network model using stochastic gradient descent; and

program instructions to use, to predict an outcome by analyzing live data, the trained neural network model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2022
From: NGUYEN, LAM MINH; TRAN, HUYEN TRANG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 060509/0185 →
Continuity (1)
Related Publication 20240020528A1 · Jan 18, 2024
References Cited (33)
US 10019470B2 · Birdwell · 2018 [cited by examiner]
US 10572800B2 · Wang et al. · 2020 [cited by applicant]
US 11003989B2 · Ho · 2021 [cited by applicant]
US 11568171B2 · Nguyen · 2023 [cited by examiner]
US 11811645B1 · Reth · 2023 [cited by examiner]
US 11947668B2 · Harang · 2024 [cited by examiner]
US 12114243B2 · Yang · 2024 [cited by examiner]
US 20170270409A1 · Trischler · 2017 [cited by examiner]
US 20190050727A1 · Anderson et al. · 2019 [cited by applicant]
US 20210073661A1 · Matlick · 2021 [cited by examiner]
US 20210232734A1 · Baayen · 2021 [cited by applicant]
US 20210294781A1 · Fernández Musoles · 2021 [cited by examiner]
US 20220171996A1 · Nguyen et al. · 2022 [cited by applicant]
US 20220198275A1 · Baker · 2022 [cited by applicant]
US 20220406404A1 · Kuznetsov · 2022 [cited by examiner]
US 20230274478A1 · Turgutlu · 2023 [cited by examiner]
Jiang; Neural network training; Univ of Rochester; 24 pages; 2022. [cited by examiner]
Jonathan; Early Phase of Neural Network Training; MIT; 20 pages; 2019. [cited by examiner]
Khomenko; Accelerating_recurrent_neural_network_training; IEEE; pp. 100-103; 2016. [cited by examiner]
Wang; Accelerating neural Network Training; Georgia IT; 12 pages; 2017. [cited by examiner]
Hu et al., Accelerated Gradient Methods for Stochastic Optimization and Online Learning, Advances in Neural Information Processing Systems, vol. 22. Curran Associates, Inc., 2009. [cited by applicant]
Lan, An optimal method for stochastic composite optimization, Mathematical Programming, 133:365-397, Jan. 1, 2011. [cited by applicant]
Mishchenko et al., Random reshuffling: Simple analysis with vast improvements, 34th Conference on Neural Information Processing Systems (NeurIPS 2020), 2020. [cited by applicant]
Nemirovski et al., Robust stochastic approximation approach to stochastic programming, SIAM J. on Optimization, vol. 19, No. 4, pp. 1574-1609, 2009. [cited by applicant]
Nesterov, A method of solving a convex programming problem with convergence rate O(1/k^2), Soviet Math. Dokl., vol. 27, No. 2, pp. 372-376, 1983. [cited by applicant]
Nesterov, Introductory lectures on convex optimization: A basic course, Applied Optimization, Kluwer Academic Publishers, vol. 87, 2004. [cited by applicant]
Nguyen et al., A unified convergence analysis for shuffling-type gradient methods, Journal of Machine Learning Research, 22 (207), pp. 1-44, Sep. 2021. [cited by applicant]
Polyak, Some methods of speeding up the convergence of iteration methods, USSR Computational Mathematics and Mathematical Physics, vol. 4, No. 5, pp. 1-17, 1964. [cited by applicant]
Robbins et al., A stochastic approximation method, The Annals of Mathematical Statistics, vol. 22, No. 3, pp. 400-407, 1951. [cited by applicant]
Shamir et al., Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes, 30th International Conference on Machine Learning, vol. 28, pp. 71-79, Jun. 17-19, 2013. [cited by applicant]
Vaswani et al., Fast and faster convergence of SGD for over-parameterized models (and an accelerated perceptron), Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics, vol. 89, pp. … [cited by applicant]
Zhong et al., Accelerated Stochastic Gradient Method for Composite Regularization, Proceedings of the 17th International Conference on Artificial Intelligence and Statistics, vol. 33, pp. 1086-1094, Apr. 22-25, 2014. [cited by applicant]
Zhou et al., SGD converges to global minimum in deep learning via star-convex path, ICLR, 2019. [cited by applicant]