IP Library Granted Patent US 11,158,303
Granted Patent B2
US 11,158,303 · App. 16/551,915 · Granted Oct 26, 2021

Soft-forgetting for connectionist temporal classification based automatic speech recognition

Inventors: Kartik Audhkhasi (White Plains, NY); George Andrei Saon (Stamford, CT); Zoltan Tueske (White Plains, NY); Brian E. D. Kingsbury (Cortlandt Manor, NY); Michael Alan Picheny (White Plains, NY)
Assignee: International Business Machines Corporation
G10L15/063G10L15/05G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,158,303
App. No.
16/551,915
Granted
Oct 26, 2021
Kind
B2
Abstract

In an approach to soft-forgetting training, one or more computer processors train a first model utilizing one or more training batches wherein each training batch of the one or more training batches comprises one or more blocks of information. The one or more computer processors, responsive to a completion of the training of the first model, initiate a training of a second model utilizing the one or more training batches. The one or more computer processors jitter a random block size for each block of information for each of the one or more training batches for the second model. The one or more computer processors unroll the second model over one or more non-overlapping contiguous jittered blocks of information. The one or more computer processors, responsive to the unrolling of the second model, reduce overfitting for the second model by applying twin regularization.

Claims (44)

1. A computer-implemented method comprising:

training, by one or more computer processors, a first model utilizing one or more training batches wherein each training batch of the one or more training batches comprises one or more acoustic sequences;

responsive to a completion of the training of the first model, initiating, by one or more computer processors, a training of a second model utilizing the one or more training batches;

jittering, by one or more computer processors, a random block size for each acoustic sequence for each of the one or more training batches for the second model;

unrolling, by one or more computer processors, the second model over one or more non-overlapping contiguous jittered acoustic sequences; and

responsive to the unrolling of the second model, reducing, by one or computer processors, overfitting for the second model by applying twin regularization.

2. The method of claim 1 , wherein the first model is a whole-utterance bidirectional long short-term memory network.

3. The method of claim 1 , wherein the second model is a chunk-based bidirectional long short-term memory network.

4. The method of claim 1 , wherein the one or more non-overlapping contiguous jittered acoustic sequences are calculated from a connectionist temporal classification loss.

5. The method of claim 1 , wherein each acoustic sequence has an associated textual label.

6. The method of claim 1 , further comprising:

responsive to a completion of the training of the second model, disposing, by one or more computer processors, the first model; and

responsive to the completion of the training of second model, deploying, by one or more computer processors, the second model to one or more production environments.

7. The method of claim 1 , wherein twin regularization comprises a loss value and a connectionist temporal classification loss of the second model, wherein the loss value is a mean-squared error between hidden states of the first model and second model and the first model.

8. A computer program product comprising:

one or more computer readable storage media and program instructions stored on the one or more computer readable storage media, the stored program instructions comprising:

program instructions to train a first model utilizing one or more training batches wherein each training batch of the one or more training batches comprises one or more acoustic sequences;

program instructions to, responsive to a completion of the training of the first model, initiate a training of a second model utilizing the one or more training batches;

program instructions to jitter a random block size for each acoustic sequence for each of the one or more training batches for the second model;

program instructions to unroll the second model over one or more non-overlapping contiguous jittered acoustic sequence; and

program instructions to, responsive to the unrolling of the second model, reduce overfitting for the second model by applying twin regularization.

9. The computer program product of claim 8 , wherein the first model is a whole-utterance bidirectional long short-term memory network.

10. The computer program product of claim 8 , wherein the second model is a chunk-based bidirectional long short-term memory network.

11. The computer program product of claim 8 , wherein the one or more non-overlapping contiguous jittered acoustic sequences are calculated from a connectionist temporal classification loss.

12. The computer program product of claim 8 , wherein the program instructions stored on the one or more computer readable storage media comprise:

program instructions to, responsive to a completion of the training of the second model, dispose, by one or more computer processors, the first model; and

program instructions to, responsive to the completion of the training of second model, deploy, by one or more computer processors, the second model to one or more production environments.

13. The computer program product of claim 8 , wherein twin regularization comprises a loss value and a connectionist temporal classification loss of the second model, wherein the loss value is a mean-squared error between hidden states of the first model and second model and the first model.

14. A computer system comprising:

one or more computer processors;

one or more computer readable storage media; and

program instructions stored on the computer readable storage media for execution by at least one of the one or more processors, the stored program instructions comprising:

program instructions to train a first model utilizing one or more training batches wherein each training batch of the one or more training batches comprises one or more acoustic sequences;

program instructions to, responsive to a completion of the training of the first model, initiate a training of a second model utilizing the one or more training batches;

program instructions to jitter a random block size for each acoustic sequence for each of the one or more training batches for the second model;

program instructions to unroll the second model over one or more non-overlapping contiguous jittered acoustic sequences; and

program instructions to, responsive to the unrolling of the second model, reduce overfitting for the second model by applying twin regularization.

15. The computer system of claim 14 , wherein the first model is a whole-utterance bidirectional long short-term memory network.

16. The computer system of claim 14 , wherein the second model is a chunk-based bidirectional long short-term memory network.

17. The computer system of claim 14 , wherein the one or more non-overlapping contiguous jittered acoustic sequences are calculated from a connectionist temporal classification loss.

18. The computer system of claim 14 , wherein the program instructions stored on the one or more computer readable storage media comprise:

program instructions to, responsive to a completion of the training of the second model, dispose, by one or more computer processors, the first model; and

program instructions to, responsive to the completion of the training of second model, deploy, by one or more computer processors, the second model to one or more production environments.

19. The computer system of claim 14 , wherein twin regularization comprises a loss value and a connectionist temporal classification loss of the second model, wherein the loss value is a mean-squared error between hidden states of the first model and second model and the first model.

Assignments (3)
SECURITY INTEREST Recorded Jul 8, 2025
From: ANTHROPIC, PBC
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 071626/0234 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 22, 2025
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: ANTHROPIC, PBC
Reel/Frame 071201/0198 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2019
From: AUDHKHASI, KARTIK; SAON, GEORGE ANDREI; TUESKE, ZOLTAN; KINGSBURY, BRIAN E. D.; PICHENY, MICHAEL ALAN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 050177/0633 →
Continuity (1)
Related Publication 20210065680A1 · Mar 4, 2021