IP Library Granted Patent US 10,380,484
Granted Patent B2
US 10,380,484 · App. 14/920,304 · Granted Aug 13, 2019

Annealed dropout training of neural networks

Inventors: Vaibhava Goel (Chappaqua, NY); Steven John Rennie (Yorktown Heights, NY); Samuel Thomas (Elmsford, NY); Ewout van den Berg (Bronxville, NY)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/082G06N3/0454G06N3/08G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,380,484
App. No.
14/920,304
Granted
Aug 13, 2019
Kind
B2
Abstract

Systems and methods for training a neural network to optimize network performance, including sampling an applied dropout rate for one or more nodes of the network to evaluate a current generalization performance of one or more training models. An optimized annealing schedule may be generated based on the sampling, wherein the optimized annealing schedule includes an altered dropout rate configured to improve a generalization performance of the network. A number of nodes of the network may be adjusted in accordance with a dropout rate specified in the optimized annealing schedule. The steps may then be iterated until the generalization performance of the network is maximized.

Claims (17)

1. A method for training a neural network to optimize network performance of a network of interconnected computers, comprising:

iteratively sampling an applied dropout rate for one or more nodes of the network to evaluate a current generalization performance of one or more training models;

iteratively generating, using a processor, an optimized annealing schedule based on the sampling, wherein the optimized annealing schedule includes an altered dropout rate configured to improve a generalization performance of the network;

increasing a realized capacity of the neural network by iteratively adjusting a number of nodes of the network in accordance with a dropout rate specified in the optimized annealing schedule until the generalization performance of the network is maximized.

2. The method as recited in claim 1 , wherein the adjusting a number of nodes further comprises zeroing a fixed percentage of the nodes.

3. The method as recited in claim 1 , wherein the adjusting a number of nodes further comprises gradually decreasing a dropout probability of the nodes in the network during the training.

4. The method as recited in claim 1 , wherein the adjusting a number of nodes further comprises increasing the dropout rate for successive iterations to prevent overfitting of training data.

5. The method as recited in claim 1 , wherein the applied dropout rate is sampled from a distribution over dropout rates, which are estimated or evolved as network training proceeds.

6. The method as recited in claim 1 , wherein the adjusting a number of nodes further comprises randomly setting the output of one or more of the nodes to zero with dropout probability p d , and wherein the adjusting a number of nodes further comprises iteratively adjusting the dropout probability p d .

7. The method as recited in claim 6 , wherein the adjusting a number of nodes further comprises applying one of a linear or geometric fixed decaying schedule to adjust the dropout probability p d .

8. The method as recited in claim 1 , wherein the generating the optimized annealing schedule further comprises generating an optimized joint annealing schedule, as indicated by a loss function on the held-out data, for at least one of a learning rate, the dropout probability p d , and any other hyperparameters of the learning procedure for two or more training models.

9. The method as recited in claim 8 , wherein generating the optimized joint annealing schedule further comprises:

considering all or a subset of the set of combinations implied by one of holding fixed, increasing, or decreasing each hyperparameter by a specified amount; and

selecting a subset, including the N best performing models, N>=1, of models that result, based on one or more iterations of learning, for application in additional training iterations.

10. The method as recited in claim 1 , wherein the training further comprises correcting biases associated with each of one or more linear projections taken at one or more network layers so that they are consistent with a subspace implied by a dropout mask applied to one or more inputs of the layer.

11. The method as recited in claim 10 , wherein the training further comprises inputting dropout mask specific biases for each of the one or more linear projections.

12. The method as recited in claim 11 , wherein the dropout mask specific biases are realized with a matrix of learned biases, and a total bias to apply to each linear projection is determined by multiplying a dropout mask vector by a matrix.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2015
From: GOEL, VAIBHAVA; RENNIE, STEVEN JOHN; THOMAS, SAMUEL; VAN DEN BERG, EWOUT
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 036858/0627 →
Continuity (3)
Continuation 14842348 · Sep 1, 2015
Provisional Application 62149624 · Apr 19, 2015
Related Publication 20160307098A1 · Oct 20, 2016
Cited By (5)
US 12,198,032 US 12,373,700 US 12,412,206 US 12,536,664 US 12,639,574