IP Library Granted Patent US 9,483,728
Granted Patent B2
US 9,483,728 · App. 14/528,638 · Granted Nov 1, 2016

Systems and methods for combining stochastic average gradient and hessian-free optimization for sequence training of deep neural networks

Inventors: Pierre Dognin (White Plains, NY); Vaibhava Goel (Chappaqua, NY)
Assignee: International Business Machines Corporation
G06N3/08G06N3/0454G06N3/084G10L15/063G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,483,728
App. No.
14/528,638
Granted
Nov 1, 2016
Kind
B2
Abstract

A method for training a deep neural network (DNN), comprises receiving and formatting speech data for the training, performing Hessian-free sequence training (HFST) on a first subset of a plurality of subsets of the speech data, and iteratively performing the HFST on successive subsets of the plurality of subsets of the speech data, wherein iteratively performing the HFST comprises reusing information from at least one previous iteration.

Claims (18)

1. A method for training a deep neural network, comprising:

receiving, using at least one processor operatively coupled to a memory of a computer system, speech data for the training;

formatting, using the at least one processor, the speech data for the training;

dividing, using the at least one processor, the speech data into a plurality of subsets;

performing, using the at least one processor, Hessian-free sequence training on a first subset of the plurality of subsets of the speech data;

iteratively performing, using the at least one processor, the Hessian-free sequence training on successive subsets of the plurality of subsets of the speech data;

wherein iteratively performing the Hessian-free sequence training comprises:

processing the first subset of the speech data to generate a first gradient of loss in a first iteration;

processing a successive subset of the speech data to generate a second gradient of loss in a second iteration;

dynamically computing weights for the first gradient of loss and for the second gradient of loss; and

reusing gradient information from at least one previous iteration, wherein reusing the gradient information from the at least one previous iteration comprises integrating a weighted first gradient of loss and a weighted second gradient of loss to generate a solution to the second iteration; and

transmitting, using the at least one processor, a result of the iterative performance of the Hessian-free sequence training to the deep neural network.

2. The method of claim 1 , wherein the gradient information comprises average gradient information.

3. The method of claim 1 , wherein the weights for first gradient of loss and the second gradient of loss are chosen based on a difference between a held-out loss of the second iteration and a held-out loss of the first iteration.

4. The method of claim 1 , wherein the dynamic computing comprises changing weights assigned to the first gradient of loss and the second gradient of loss for different iterations.

5. The method of claim 1 , wherein the dynamic computing comprises dynamically estimating weights assigned to the first gradient of loss and the second gradient of loss to provide a loss function gradient before an iteration takes place.

6. The method of claim 1 , wherein weights assigned to the first gradient of loss and the second gradient of loss are based on a tunable parameter that controls exponentiation of the weights across the plurality of subsets.

7. The method of claim 1 , wherein the dynamic computing comprises using a weighting function to assign higher weights to gradients with held-out loss values closer to a held-out loss value of a current model, and lower weights to gradients with held-out loss values farther from the held-out loss value of the current model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2014
From: DOGNIN, PIERRE; GOEL, VAIBHAVA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 034073/0878 →
Continuity (2)
Provisional Application 61912638 · Dec 6, 2013
Related Publication 20150161988A1 · Jun 11, 2015