IP Library › Granted Patent US 12,619,879
Granted Patent B2
US 12,619,879 · App. 17/481,160 · Granted May 5, 2026

Training neural networks using learned optimizers

Inventors: Luke Shekerjian Metz (Mountain View, CA); Niruban Maheswaranathan (San Jose, CA); Christian Daniel Freeman (San Francisco, CA); Benjamin Poole (Palo Alto, CA); Jascha Narain Sohl-Dickstein (San Francisco, CA)
Assignee: Google LLC
G06N3/086
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,879
App. No.
17/481,160
Granted
May 5, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a neural network. One of the methods includes performing, using a plurality of training examples, a training step to obtain respective gradients of a loss function with respect to each of the parameters in the parameter tensors; obtaining a validation loss for a plurality of validation examples that are different from the plurality of training examples generating an optimizer input from at least the respective gradients and the validation loss; processing the optimizer input using an optimizer neural network to generate an output defining a respective update for each of the parameters in the parameter tensors of the neural network; and for each of the parameters in the parameter tensors, applying the respective update to a current value of the parameter to generate an updated value for the parameter.

Claims (42)

1 . A method of training a first neural network configured to perform a machine learning task by processing a network input in accordance with at least a set of parameter tensors each including a plurality of respective parameters to generate a network output for the machine learning task, the method comprising repeatedly performing operations comprising:

performing, using a plurality of training examples, a training step to obtain respective gradients of a loss function for the machine learning task with respect to each of the parameters in the parameter tensors of the first neural network;

obtaining a validation loss that measures a performance of the first neural network on the machine learning task for a plurality of validation examples that are different from the plurality of training examples used to obtain the respective gradients;

processing an optimizer input that includes i) features derived from the respective gradients of the loss function for the machine learning task obtained using the training examples and ii) features that are computed from the validation loss that measures a performance of the first neural network on the plurality of validation examples that are different from the plurality of training examples using an optimizer neural network to generate an output defining a respective update for each of the parameters in the parameter tensors of the first neural network, wherein the respective updates automatically regularize the training of the first neural network and reduce time required to train the first neural network; and

for each of the parameters in the parameter tensors, applying the respective update to a current value of the parameter to generate an updated value for the parameter.

2 . The method of claim 1 , further comprising:

generating, from results of the training step, training data for training the optimizer neural network; and

performing a training step to train the optimizer neural network on the training data to optimize an objective that measures a performance of the optimizer neural network in generating at least the respective updates.

3 . The method of claim 2 , wherein the objective measures (i) the performance of the optimizer neural network in generating the respective updates and (ii) a performance of the optimizer neural network in generating updates during training a plurality of other neural networks to perform a plurality of other machine learning tasks.

4 . The method of claim 2 , wherein performing the training step comprises performing one or more iterations of an evolution strategies (ES) technique to optimize the objective.

5 . The method of claim 1 , wherein the optimizer neural network has been trained to optimize an objective that measures a quality of parameter updates generated by the optimizer neural network for a plurality of machine learning tasks that does not include the machine learning task.

6 . The method of claim 1 , wherein the optimizer neural network comprises:

(i) a per-tensor neural network that operates independently for each of the parameter tensors, and

(ii) a per-parameter neural network that operates independently for each of the plurality of parameters of each of the parameter tensors.

7 . The method of claim 6 , wherein the per-tensor neural network is a recurrent neural network and the per-parameter neural network is a feedforward neural network.

8 . The method of claim 7 , wherein the per-parameter neural network is a multi-layer perceptron.

9 . The method of claim 6 , wherein the per-parameter neural network generates, for each parameter, an output that comprises (i) a direction for the parameter update for the parameter and (ii) a magnitude value for the parameter update for the parameter.

10 . The method of claim 9 , further comprising:

for each parameter, generating the update, comprising exponentiating the magnitude value for the parameter to generate an exponentiation and multiplying the exponentiation by the direction for the parameter to generate a product.

11 . The method of claim 10 , wherein generating the update further comprises applying gradient clipping to the product to generate the update.

12 . The method of claim 6 , wherein the optimizer input comprises a respective tensor input for each of the parameter tensors and a respective parameter input for each of the parameters of each of the parameter tensors.

13 . The method of claim 12 , wherein generating the optimizer input comprises:

generating the tensor input for each of the parameter tensors from at least (i) the validation loss and (ii) gradients for the parameters in the parameter tensor.

14 . The method of claim 13 , wherein generating the tensor input for each of the parameter tensors further comprises generating the tensor input from at least (iii) a training loss for the training step for the plurality of training examples.

15 . The method of claim 12 , wherein generating the tensor input for each of the parameter tensors further comprises generating the tensor input from at least (iv) outputs generated by the per-tensor neural network when updating the parameters at a preceding training step.

16 . The method of claim 12 wherein generating the tensor input for each of the parameter tensors further comprises generating the tensor input from at least (v) outputs generated by the per-parameter neural network for the parameters in the corresponding parameter tensor when updating the parameters at a preceding training step.

17 . The method of claim 12 , wherein generating the optimizer input comprises:

generating the parameter input for each of the parameters from at least (i) the gradient for the parameter and (ii) an output of the per-tensor neural network generated by processing the corresponding tensor input for the parameter tensor to which the parameter belongs.

18 . The method of claim 17 , wherein generating the parameter input for each of the parameters further comprises generating the parameter input from at least (iii) a current value of the parameter.

19 . The method of claim 1 , wherein:

the optimizer neural network generates updates to the parameters at each training step in a sequence of training steps, and

the validation loss is updated after each of a proper subset of the training steps.

20 . One or more non-transitory computer-readable media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for training a first neural network configured to perform a machine learning task by processing a network input in accordance with at least a set of parameter tensors each including a plurality of respective parameters to generate a network output for the machine learning task, the operations comprising:

performing, using a plurality of training examples, a training step to obtain respective gradients of a loss function for the machine learning task with respect to each of the parameters in the parameter tensors of the first neural network;

obtaining a validation loss that measures a performance of the first neural network on the machine learning task for a plurality of validation examples that are different from the plurality of training examples used to obtain the respective gradients;

processing an optimizer input that includes i) features derived from the respective gradients of the loss function for the machine learning task obtained using the training examples and ii) features that are computed from the validation loss that measures a performance of the first neural network on the plurality of validation examples that are different from the plurality of training examples using an optimizer neural network to generate an output defining a respective update for each of the parameters in the parameter tensors of the first neural network, wherein the respective updates automatically regularize the training of the first neural network and reduce time required to train the first neural network; and

for each of the parameters in the parameter tensors, applying the respective update to a current value of the parameter to generate an updated value for the parameter.

21 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by one or more computers cause the one or more computers to perform operations for training a first neural network configured to perform a machine learning task by processing a network input in accordance with at least a set of parameter tensors each including a plurality of respective parameters to generate a network output for the machine learning task, the operations comprising:

performing, using a plurality of training examples, a training step to obtain respective gradients of a loss function for the machine learning task with respect to each of the parameters in the parameter tensors of the first neural network;

obtaining a validation loss that measures a performance of the first neural network on the machine learning task for a plurality of validation examples that are different from the plurality of training examples used to obtain the respective gradients;

processing an optimizer input that includes i) features derived from the respective gradients of the loss function for the machine learning task obtained using the training examples and ii) features that are computed from the validation loss that measures a performance of the first neural network on the plurality of validation examples that are different from the plurality of training examples using an optimizer neural network to generate an output defining a respective update for each of the parameters in the parameter tensors of the first neural network, wherein the respective updates automatically regularize the training of the first neural network and reduce time required to train the first neural network; and

for each of the parameters in the parameter tensors, applying the respective update to a current value of the parameter to generate an updated value for the parameter.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2021
From: METZ, LUKE SHEKERJIAN; MAHESWARANATHAN, NIRUBAN; FREEMAN, CHRISTIAN DANIEL; POOLE, BENJAMIN; SOHL-DICKSTEIN, JASCHA NARAIN
To: GOOGLE LLC
Reel/Frame 057803/0579 →
Continuity (2)
Provisional Application 63081269 · Sep 21, 2020
Related Publication 20220092429A1 · Mar 24, 2022
References Cited (81)
US 11941527B2 · Jaderberg · 2024 [cited by examiner]
US 12112254B1 · Teig · 2024 [cited by examiner]
US 20020174079A1 · Mathias · 2002 [cited by examiner]
US 20180240010A1 · Faivishevsky · 2018 [cited by examiner]
US 20200380369A1 · Case · 2020 [cited by examiner]
US 20210192357A1 · Sinha · 2021 [cited by examiner]
Auther: Wichrowska Title: Learned Optimizers that Scale and Generalize: (Year: 2017). [cited by examiner]
Abadi et al., “Tensorflow: A system for large-scale machine learning,” Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation, Nov. 2016, 16:265-283. [cited by applicant]
Andrychowicz et al., “Learning to learn by gradient descent by gradient descent,” Advances in Neural Information Processing Systems 29, Dec. 2016, 9 pages. [cited by applicant]
Baydin et al., “Automatic differentiation of algorithms for machine learning,” arXiv, Apr. 28, 2014, 7 pages. [cited by applicant]
Bello et al., “Neural optimizer search with reinforcement learning,” arXiv, Sep. 22, 2017, 12 pages. [cited by applicant]
Bengio et al., “Advances in optimizing recurrent networks,” Presented at IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada, May 26-31, 2013, pp. 8624-8628. [cited by applicant]
Bengio, “Gradient-based optimization of hyperparameters,” Neural Computation, Jan. 2000, 12:1890-1900. [cited by applicant]
Berner et al., “Dota 2 with large scale deep reinforcement learning,” arXiv, Dec. 13, 2019, 66 pages. [cited by applicant]
Choi et al., “On empirical comparisons of optimizers for deep learning,” arXiv, Oct. 11, 2019, 29 pages. [cited by applicant]
Choromanski et al., “Structured evolution with compact architectures for scalable policy optimization,” arXiv, Jun. 12, 2018, 16 pages. [cited by applicant]
Chung et al., “Empirical evaluation of gated recurrent neural networks on sequence modeling,” arXiv, Dec. 11, 2014. [cited by applicant]
Cobbe et al., “Quantifying generalization in reinforcement learning,” arXiv, Dec. 20, 2018, 20 pages. [cited by applicant]
cs.toronto.edu [online], “Cifar-10 and cifar-100 datasets,” Sep. 23, 2009, retrieved on Feb. 15, 2022, retrieved from URL<https://www.cs.toronto.edu/˜kriz/cifar.html>, 4 pages. [cited by applicant]
Daniel et al., “Learning step size controllers for robust neural network training,” Proceedings of the AAAI Conference on Artificial Intelligence, Feb. 21, 2016, 30(1):1519-1525. [cited by applicant]
DeepMind.com [online], “Alphastar: Mastering the real-time strategy game StarCraft II,” May 20, 2000, retrieved on Jan. 24, 2019, retrieved from URL<https://deepmind.com/blog/article/alphastar-mastering-real-time-strate… [cited by applicant]
Devlin et al., “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv, Oct. 11, 2018, 14 pages. [cited by applicant]
Dozat, “Incorporating Nesterov Momentum into Adam,” Computer Science, Feb. 18, 2016, 4 pages. [cited by applicant]
Finn et al., “Model-agnostic meta-learning for fast adaptation of deep networks,” arXiv, Jul. 18, 2017, 13 pages. [cited by applicant]
Franceschi et al., “Bilevel programming for hyperparameter optimization and meta-learning,” arXiv, Jul. 3, 2018, 12 pages. [cited by applicant]
Ghemawat et al., “The google file system,” Proceedings of the nineteenth ACM symposium on Operating systems principles, Dec. 2003, pp. 29-43. [cited by applicant]
Golovin et al., “Google Vizier: A Service for Black-Box Optimization,” Presented at International Conference on Knowledge Discovery and Data Mining, Halifax, NS, Canada, Aug. 13-17, 2017, 10 pages. [cited by applicant]
Graves et al., “Neural Turing Machines,” arXiv, Dec. 10, 2014, 26 pages. [cited by applicant]
Gu et al., “Meta-learning biologically plausible semi-supervised update rules,” bioRxiv, Dec. 30, 2019, 6 pages. [cited by applicant]
Harris et al., “Array programming with NumPy,” Nature, Sep. 17, 2020, 585:357-362. [cited by applicant]
Hart et al., “The new compiler,” AI Memo 39, 1962, 6 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2016, 9 pages. [cited by applicant]
He et al., “Identity Mappings in Deep Residual Networks,” European Conference on Computer Vision, Oct. 2016, 16 pages. [cited by applicant]
Heess et al., “Emergence of Locomotion Behaviours in Rich Environments,” arXiv, Jul. 10, 2017, 14 pages. [cited by applicant]
Hochreiter et al., “Long Short-Term Memory,” Neural Computation, 1997, 32 pages. [cited by applicant]
Hofmann et al., “Letter-Value Plots: Boxplots for Large Data,” Technical Report, 2011, 22 pages. [cited by applicant]
Hubinger et al., “Risks from learned optimization in advanced machine learning systems,” arXiv, Jun. 11, 2019, 39 pages. [cited by applicant]
Hunter, “Matplotlib: A 2D graphics environment,” Computing in Science & Engineering, Jun. 2007, 9(3):90-95. [cited by applicant]
Kingma et al., “Adam: A method for stochastic optimization,” arXiv, Dec. 22, 2014, 9 pages. [cited by applicant]
Kingma et al., “Auto-encoding variational bayes,” arXiv, Dec. 27, 2013, 14 pages. [cited by applicant]
Krizhevsky et al., “ImageNet classification with deep convolutional neural networks,” Advances in Neural Information Processing Systems 25, Dec. 2012, 9 pages. [cited by applicant]
Li et al., “Learning to Optimize Neural Nets,” arXiv, Nov. 30, 2017, 10 pages. [cited by applicant]
Li et al., “Learning to Optimize,” Presented at International Conference on Learning Representations, Toulon, France, Apr. 24-26, 2017, 13 pages. [cited by applicant]
Li et al., “Meta SGD-Learning to Learn Quickly for Few-Shot Learning,” arXiv print: 1707.0983, Sep. 28, 2017, 11 pages. [cited by applicant]
Liu et al., “On the Limited Memory BFGS Method for Large Scale Optimization,” Mathematical Programming 45, Aug. 1989, 27 pages. [cited by applicant]
Loshchilov et al., “Decoupled weight decay regularization,” arXiv, Jan. 4, 2019, 19 pages. [cited by applicant]
Loshchilov et al., “SGDR: Stochastic gradient descent with warm restarts,” arXiv, May 3, 2017, 16 pages. [cited by applicant]
Lv et al., “Learning gradient descent: Better generalization and longer horizons,” arXiv, Jun. 10, 2017, 9 pages. [cited by applicant]
Maclaurin et al., “Gradient-based Hyperparameter Optimization through Reversible Learning,” Presented at International Conference on Machine Learning, Lille, France, Jul. 6-11, 2015, 10 pages. [cited by applicant]
Maheswaranathan et al., “Guided evolutionary strategies: Augmenting random search with surrogate gradients,” Presented at International Conference on Machine Learning, Long Beach, CA, Jun. 9-15, 2019, 10 pages. [cited by applicant]
Metz et al., “Learning unsupervised learning rules,” arXiv, May 23, 2018, 25 pages. [cited by applicant]
Metz et al., “Understanding and correcting pathologies in the training of learned optimizers,” Presented at International Conference on Machine Learning, Long Beach, CA, Jun. 9-15, 2019, 10 pages. [cited by applicant]
Metz et al., “Using a thousand optimization tasks to learn hyperparameter search strategies,” arXiv, Apr. 1, 2020, 34 pages. [cited by applicant]
Metz et al., “Using learned optimizers to make models robust to input noise,”, arXiv, Jun. 8, 2019, 8 pages. [cited by applicant]
Nesterov et al., “Random gradient-free minimization of convex functions,” Core Discussion Paper, Jan. 2011, 34 pages. [cited by applicant]
Nichol et al., “On first-order meta-learning algorithms,” arXiv, Oct. 22, 2018, 15 pages. [cited by applicant]
Nocedal, “Updating Quasi-Newton Matrices with Limited Storage,” Mathematics of Computation, Jul. 1980, 35(151):773-782. [cited by applicant]
Papamakarios et al., “Masked autoregressive flow for density estimation,” Advances in Neural Information Processing Systems 30, Dec. 2017, 10 pages. [cited by applicant]
Pearlmutter, “An investigation of the gradient descent process in neural networks, ” Thesis for the degree of Doctor of Philosophy, Carnegie Mellon University, School of Computer Science, Jan. 16, 1996, 140 pages. [cited by applicant]
Peters et al., “Relative Entropy Policy Search,” Proceedings of the Twenty-Fourth AAAI Conference on Artificial Intelligence, Jul. 2010, pp. 1607-1612. [cited by applicant]
Piech et al., “Deep knowledge tracing”, Advances in neural information processing systems 28, Dec. 2015, 9 pages. [cited by applicant]
Rastrigin, “About convergence of random search method in extremal control of multiparameter systems,” Avtomat. i Telemekh, 1963, 24(11):1467-1473 (with English abstract on the last page). [cited by applicant]
Runarsson et al., “Evolution and Design of Distributed Learning Rules,” Presented at the IEEE Symposium on Combinations of Evolutionary Computation and Neural Networks, San Antonio, TX, May 11-13, 2000, 5 pages. [cited by applicant]
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision, Apr. 11, 2015, 42 pages. [cited by applicant]
Salimans et al., “Evolution strategies as a scalable alternative to reinforcement learning,” arXiv, Sep. 7, 2017, 13 pages. [cited by applicant]
Schulman et al., “Proximal Policy Optimization Algorithms,” arXiv, Aug. 28, 2017, 12 pages. [cited by applicant]
Shahriari et al., “Taking the human out of the loop: A review of bayesian optimization.,” Proceedings of the IEEE, Jan. 2016, 104(1):148-175. [cited by applicant]
Shazeer et al., “Adafactor: Adaptive Learning Rates with Sublinear Memory Cost,” arXiv, Apr. 11, 2018, 9 pages. [cited by applicant]
Sivaprasad et al., “On the Tunability of Optimizers in Deep Learning,” arXiv, Oct. 25, 2019, 16 pages. [cited by applicant]
Staines et al., “Variational Optimization,” arXiv print: 1212.4507, Dec. 20, 2012, 14 pages. [cited by applicant]
Strubell et al., “Energy and policy considerations for deep learning in NLP,” arXiv, Jun. 5, 2019, 6 pages. [cited by applicant]
Tallec et al., “Unbiasing Truncated Backpropagation Through Time,” arXiv, May 23, 2017, 13 pages. [cited by applicant]
Van Der Walt et al., “The NumPy array: A structure for efficient numerical computation,” Computing in Science & Engineering, Mar./Apr. 2011, 13(2):22-30. [cited by applicant]
Vaswani et al., “Attention Is All You Need,” Advances in Neural Information Processing Systems 30, Dec. 2017, 11 pages. [cited by applicant]
Werbos, “Backpropagation Through Time: What It Does and How to Do It,” Proceedings of the IEEE, Oct. 1990, 78(10):1550-1560. [cited by applicant]
Wichrowska et al., “Learned optimizers that scale and generalize,” Presented at International Conference on Machine Learning, Sydney-Australia, Aug. 6-11, 2017, 10 pages. [cited by applicant]
Xu et al., “Learning an Adaptive Learning Rate Schedule,” arXiv, Sep. 20, 2019, 6 pages. [cited by applicant]
Xu et al., “Reinforcement Learning for Learning Rate Control,” arXiv, May 31, 2017, 7 pages. [cited by applicant]
yan.lecun.com [online], “The mnist database of handwritten digits,” 1998, 5 pages. [cited by applicant]
zenodo.org [online], “mwaskom/seaborn: v0.11.0 (Sep. 2020),” Sep. 8, 2020, retrieved on Feb. 16, 2022, retrieved from URL <https://zenodo.org/record/4019146#.Yg0hsYjMKUk>, 5 pages. [cited by applicant]
Zoph et al., “Neural Architecture Search with Reinforcement Learning,” arXiv print: 1611.01578, Feb. 15, 2017, 16 pages. [cited by applicant]