IP Library Granted Patent US 12,271,823
Granted Patent B2
US 12,271,823 · App. 18/180,754 · Granted Apr 8, 2025

Training machine learning models by determining update rules using neural networks

Inventors: Misha Man Ray Denil (London, GB); Tom Schaul (London, GB); Marcin Andrychowicz (London, GB); Joao Ferdinando Gomes de Freitas (London, GB); Sergio Gomez Colmenarejo (London, GB); Matthew William Hoffman (London, GB); David Benjamin Pfau (London, GB)
Assignee: DeepMind Technologies Limited
G06N3/084G06N3/044G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,271,823
App. No.
18/180,754
Granted
Apr 8, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media for training machine learning models. One method includes obtaining a machine learning model, wherein the machine learning model comprises one or more model parameters, and the machine learning model is trained using gradient descent techniques to optimize an objective function; determining an update rule for the model parameters using a recurrent neural network (RNN); and applying a determined update rule for a final time step in a sequence of multiple time steps to the model parameters.

Claims (34)

1. A computer-implemented method for training a target machine learning model having a plurality of model parameters using an optimizer neural network having a plurality of optimizer parameters, the method comprising, at each current iteration of a plurality of iterations:

determining a respective update rule for each of the plurality of model parameters using the optimizer neural network by operating the optimizer neural network independently on each of the plurality of model parameters of the target machine learning model, the determining comprising, for each respective model parameter:

generating a respective parameter-specific input that comprises a gradient of a target objective function of the target machine learning model with respect to the respective model parameter; and

processing the respective parameter-specific input using the optimizer neural network and in accordance with current values of the optimizer parameters to generate a respective optimizer output that specifies a respective parameter-specific update rule for updating the respective model parameter;

applying the update rules generated by the optimizer neural network to the model parameters of the target machine learning model to update values of the model parameters; and

updating the current values of the optimizer parameters by using gradient descent techniques to minimize an optimizer objective function that depends at least on a function value of the target objective function of the target machine learning model computed using the values of the model parameters of the target machine learning model that have been updated at the current iteration.

2. The method of claim 1 , wherein applying the determined update rule for a final iteration in the plurality of iterations to the model parameters generates trained model parameters.

3. The method of claim 1 , wherein the target machine learning model comprises a neural network.

4. The method of claim 1 , wherein the current iteration is an iteration after the first iteration of the plurality of iterations, and the optimizer objective function further depends on the values of the model parameters at one or more iterations that precede the current iteration.

5. The method of claim 1 , wherein the optimizer neural network is a recurrent neural network (RNN).

6. The method of claim 5 , further comprising, at each of the plurality of iterations, providing a previous hidden state of the RNN as input to the RNN for the iteration.

7. The method of claim 5 , wherein the optimizer neural network is a long short-term memory (LSTM) neural network.

8. The method of claim 1 , further comprising preprocessing the respective parameter-specific inputs to the optimizer neural network to disregard gradients that are smaller than a predetermined threshold.

9. The method of claim 1 , wherein the optimizer objective function further depends on function values computed by applying the target objective function of the target machine learning model to the values of the model parameters that have been computed at one or more iterations that precede the current iteration.

10. The method of claim 9 , wherein the optimizer objective function is computed at the current iteration based on a weighted sum of function values of the target objective function of the target machine learning model, wherein each function value is computed by applying the target objective function to the values of the model parameters that have been computed at the current iteration or at an iteration that precede the current iteration.

11. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations for training a target machine learning model having a plurality of model parameters using an optimizer neural network having a plurality of optimizer parameters, the operations comprising, at each current iteration of a plurality of iterations:

determining a respective update rule for each of the plurality of model parameters using the optimizer neural network by operating the optimizer neural network independently on each of the plurality of model parameters of the target machine learning model, the determining comprising, for each respective model parameter:

generating a respective parameter-specific input that comprises a gradient of a target objective function of the target machine learning model with respect to the respective model parameter; and

processing the respective parameter-specific input using the optimizer neural network and in accordance with current values of the optimizer parameters to generate a respective optimizer output that specifies a respective parameter-specific update rule for updating the respective model parameter;

applying the update rules generated by the optimizer neural network to the model parameters of the target machine learning model to update values of the model parameters; and

updating the current values of the optimizer parameters by using gradient descent techniques to minimize an optimizer objective function that depends at least on a function value of the target objective function of the target machine learning model computed using the values of the model parameters that have been updated at the current iteration.

12. The system of claim 11 , wherein applying the determined update rule for a final iteration in the plurality of iterations to the model parameters generates trained model parameters.

13. The system of claim 11 , wherein the target machine learning model comprises a neural network.

14. The system of claim 11 , wherein the optimizer objective function further depends on the values of the model parameters at one or more iterations that precede the current iteration.

15. The system of claim 11 , wherein the optimizer neural network is a recurrent neural network (RNN).

16. One or more non-transitory computer-readable storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations for training a target machine learning model having a plurality of model parameters using an optimizer neural network having a plurality of optimizer parameters, the operations comprising, at each current iteration of a plurality of iterations:

determining a respective update rule for each of the plurality of model parameters using the optimizer neural network by operating the optimizer neural network independently on each of the plurality of model parameters of the target machine learning model, the determining comprising, for each respective model parameter:

generating a respective parameter-specific input that comprises a gradient of a target objective function of the target machine learning model with respect to the respective model parameter; and

processing the respective parameter-specific input using the optimizer neural network and in accordance with current values of the optimizer parameters to generate a respective optimizer output that specifies a respective parameter-specific update rule for updating the respective model parameter;

applying the update rules generated by the optimizer neural network to the model parameters of the target machine learning model to update values of the model parameters; and

updating the current values of the optimizer parameters by using gradient descent techniques to minimize an optimizer objective function that depends at least on a function value of the target objective function of the target machine learning model computed using the values of the model parameters that have been updated at the current iteration.

17. The one or more non-transitory computer-readable storage media of claim 16 , wherein applying the determined update rule for a final iteration in the plurality of iterations to the model parameters generates trained model parameters.

18. The one or more non-transitory computer-readable storage media of claim 16 , wherein the target machine learning model comprises a neural network.

19. The one or more non-transitory computer-readable storage media of claim 16 , wherein the optimizer objective function further depends on the values of the model parameters at one or more iterations that precede the current iteration.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE ASSIGNOR NAME PREVIOUSLY RECORDED AT REEL: 063950 FRAME: 0836. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 4, 2023
From: GOOGLE LLC
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 064498/0480 →
ENTITY CONVERSION Recorded Aug 4, 2023
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 064498/0630 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2023
From: DENIL, MISHA MAN RAY; SCHAUL, TOM; ANDRYCHOWICZ, MARCIN; GOMES DE FREITAS, JOAO FERDINANDO; COLMENAREJO, SERGIO GOMEZ; HOFFMAN, MATTHEW WILLIAM; PFAU, DAVID BENJAMIN
To: GOOGLE INC.
Reel/Frame 063950/0721 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2023
From: GOOGLE INC.
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 063950/0836 →