IP Library Granted Patent US 11,650,968
Granted Patent B2
US 11,650,968 · App. 16/707,265 · Granted May 16, 2023

Systems and methods for predictive early stopping in neural network training

Inventors: Dhruv Nair (New York, NY); Gideon Mendels (New York, NY); Nimrod Lahav (Hoboken, NJ)
G06F16/2246G06N3/084G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,650,968
App. No.
16/707,265
Granted
May 16, 2023
Kind
B2
Abstract

Systems and methods may train neural networks (NNs) and determine when to stop training to not waste computing or other resources when improvement is not no longer likely. After training period for a NN, a model trained using training data from other NNs may return a a probability of improvement in the loss of the NN or a probability that the likely best loss of the NN is lower than the best loss of the other NNs for which hyperparameters have been chosen. Training may be stopped if the probability is less than a threshold, or a wait value is greater than a wait threshold.

Claims (28)

1. A method of training a neural network (NN), the method comprising:

over a series of NN training epochs, where in each epoch the NN undergoes training and a loss is computed:

determining, using:

a set of model parameters;

data describing training loss of the NN; and

a model that has been trained using training losses of a plurality of NNs other than the NN;

a probability of improvement in the loss of the NN; and

if the probability is less than a threshold, or a wait value is greater than a wait threshold, stopping training.

2. The method of claim 1 , comprising if the probability is not less than a threshold and a wait value is greater than a wait threshold, continuing training.

3. The method of claim 1 , comprising increasing the wait value if the current loss of the NN is not less than the minimum loss in the loss history for the NN.

4. The method of claim 1 , comprising setting the wait value to zero if the current loss of the NN is less than or equal to the minimum loss in the loss history for the NN.

5. The method of claim 1 , wherein determining using a model a probability of improvement in the loss of the NN comprises obtaining from a leaf of at least one tree data structure data relevant to the NN.

6. The method of claim 1 , wherein determining an expected probability of improvement comprises determining a mean expected training loss and a variance.

7. The method of claim 1 , wherein model parameters comprise hyperparameters.

8. A system of training a neural network (NN), the system comprising:

a memory; and

a processor configured to:

over a series of NN training epochs, where in each epoch the NN undergoes training and a loss is computed:

determine, using:

a set of model parameters;

data describing training loss of the NN; and

a model that has been trained using training losses of a plurality of NNs other than the NN;

a probability of improvement in the loss of the NN; and

if the probability is less than a threshold, or a wait value is greater than a wait threshold, determine to stop training.

9. The system of claim 8 , wherein the processor is configured to, if the probability is not less than a threshold and a wait value is greater than a wait threshold, determine to continue training.

10. The system of claim 8 , wherein the processor is configured to increase the wait value if the current loss of the NN is not less than the minimum loss in the loss history for the NN.

11. The system of claim 8 , wherein the processor is configured to set the wait value to zero if the current loss of the NN is less than or equal to the minimum loss in the loss history for the NN.

12. The system of claim 8 , wherein determining using a model a probability of improvement in the loss of the NN comprises obtaining from a leaf of at least one tree data structure data relevant to the NN.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2020
From: NAIR, DHRUV; MENDELS, GIDEON; LAHAV, NIMROD
To: COMET ML, INC.
Reel/Frame 054015/0222 →
Continuity (2)
Provisional Application 62852525 · May 24, 2019
Related Publication 20200372342A1 · Nov 26, 2020
Cited By (5)
US 12,205,022 US 12,288,013 US 12,646,503 US 12,657,468 US 12,710,981