IP Library › Granted Patent US 9,824,683
Granted Patent B2
US 9,824,683 · App. 14/977,674 · Granted Nov 21, 2017

Data augmentation method based on stochastic feature mapping for automatic speech recognition

Inventors: Xiaodong Cui (White Plains, NY); Vaibhava Goel (Chappaqua, NY); Brian E. D. Kingsbury (Cortlandt Manor, NY)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G10L15/063G10L15/02G10L15/16G10L21/0272
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,824,683
App. No.
14/977,674
Granted
Nov 21, 2017
Kind
B2
Abstract

A method of augmenting training data includes converting a feature sequence of a source speaker determined from a plurality of utterances within a transcript to a feature sequence of a target speaker under the same transcript, training a speaker-dependent acoustic model for the target speaker for corresponding speaker-specific acoustic characteristics, estimating a mapping function between the feature sequence of the source speaker and the speaker-dependent acoustic model of the target speaker, and mapping each utterance from each speaker in a training set using the mapping function to multiple selected target speakers in the training set.

Claims (14)

1. A computer program product for augmenting training data, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to:

receive a training set of sampled audio data including speech of a plurality of source speakers;

for each of the plurality of source speakers:

convert a feature sequence of a source speaker, of the plurality of source speakers, determined from a plurality of utterances of scripted speech within the training set, to a feature sequence of a respective target speaker under the same scripted speech, wherein the feature sequence of the respective target speaker is added to the training set;

train a speaker-dependent acoustic model for the respective target speaker for corresponding speaker-specific acoustic characteristics; and

estimate a mapping function between the feature sequence of the source speaker and the speaker-dependent acoustic model of the respective target speaker; and

for each of the plurality of source speakers:

map each of the utterances from each of the plurality of source speakers in the training set using the mapping function to a plurality of other source speakers of the plurality of source speakers, wherein the mapping is added to the training set to generate augmented training data configured to train an automatic system recognition computer system.

2. The computer program product of claim 1 , wherein the speaker-dependent acoustic model is estimated using a criteria including one or more of a maximum likelihood criterion, a maximum mutual information criterion and a minimum phone error criterion.

3. The computer program product of claim 1 , wherein the mapping function is one of a linear function, an affine function and a nonlinear function.

4. The computer program product of claim 1 , further comprising a program instructions executable by the processor to cause the processor to train the automatic system recognition computer system using the augmented training data.

5. The computer program product of claim 1 , further comprising a program instructions executable by the processor to cause the processor to select the plurality of other source speakers randomly.

6. The computer program product of claim 1 , further comprising a program instructions executable by the processor to cause the processor to select the plurality of other source speakers using a criterion including at least one of vocal tract length, dialect, and signal-to-noise ratio.

7. The computer program product of claim 1 , wherein the respective target speaker is a new speaker created by perturbing a vocal tract length of the source speaker.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2015
From: CUI, XIAODONG; GOEL, VAIBHAVA; KINGSBURY, BRIAN E. D.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 037346/0061 →
Continuity (2)
Continuation 14689730 · Apr 17, 2015
Related Publication 20170200446A1 · Jul 13, 2017