IP Library › Granted Patent US 7,072,834
Granted Patent B2
US 7,072,834 · App. 10/115,934 · Granted Jul 4, 2006

Adapting to adverse acoustic environment in speech processing using playback training data

Assignee: Intel Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,072,834
App. No.
10/115,934
Granted
Jul 4, 2006
Kind
B2
Abstract

An arrangement is provided for an automatic speech recognition mechanism to adapt to an adverse acoustic environment. Some of the original training data, collected from an original acoustic environment, is played back in an adverse acoustic environment. The playback data is recorded in the adverse acoustic environment to generate recorded playback data. An existing speech model is then adapted with respect to the adverse acoustic environment based on the recorded playback data and/or the original training data.

Claims (61)

1. A method for adapting a speech processing system with speech models trained in a first acoustic environment to a second acoustic environment, comprising:

selecting at least a portion of original training data and playing back the selected data in the second acoustic environment to generate playback data, when the speech processing system is used in the second environment, the original training data being collected in the first acoustic environment and used to train the speech models;

recording the playback data in the second acoustic environment to generate recorded playback data; and

adapting the speech models from the first acoustic environment to the second acoustic environment based at least in part on the recorded playback data.

2. The method according to claim 1 , wherein adapting the speech models comprises re-training at least one of the speech models using at least the recorded playback data.

3. The method according to claim 1 , wherein the speech models comprise at least one of:

an acoustic model for describing the acoustic realization of at least one speech sound;

a background model for describing an acoustic environment; or

a transformation for altering some property of a given function when the transformation is applied to the function.

4. The method according to claim 1 , wherein adapting the speech models comprises:

estimating discrepancy between the original training data and the recorded playback data; and

adapting at least one of the speech models from the first acoustic environment to the second acoustic environment based on the discrepancy.

5. The method according to claim 1 , further comprising:

receiving input speech collected in the second acoustic environment; and

processing the input speech using the adapted speech models.

6. The method according to claim 4 , wherein adapting the speech models comprises:

generating a background model based on the original training data and the recorded playback data; and

producing a second acoustic model for the second acoustic environment based on the generated background model and a first acoustic model, the first acoustic model being generated using the original training data.

7. The method according to claim 4 , wherein adapting the speech models comprises:

deriving a transformation based on the speech models and the discrepancy; and

applying the transformation to the speech models to produce the adapted speech models.

8. A system for adapting a speech processing system with speech models trained in a first acoustic environment to a second acoustic environment, comprising:

a training data sampling and playback mechanism for selecting at least a portion of original training data and for playing back the selected data in the second acoustic environment to generate playback data when the speech processing system is used in the second environment, the original training data being collected in the first acoustic environment and used to train the speech models; and

an adverse acoustic environment adaptation mechanism for adapting the speech models from the first acoustic environment to the second environment based at least in part on the playback data.

9. The system according to claim 8 , wherein the adverse acoustic environment adaptation mechanism comprises:

a speech model re-training mechanism for re-training at least one of the speech models based on at least one of the playback data or the original training data to derive the adapted speech models for the second acoustic environment.

10. The system according to claim 8 , wherein the speech processing system comprises an automatic speech recognition mechanism for processing input speech using the adapted speech models.

11. The system according to claim 9 , wherein the speech model re-training mechanism comprises at least one of:

an acoustic model re-training mechanism for re-training an acoustic model to adapt to the second acoustic environment based on the playback data; or

a background model re-training mechanism for re-training a background model to adapt to the second acoustic environment based on the playback data.

12. The system according to claim 8 , wherein

the adverse acoustic environment adaptation mechanism comprises a discrepancy based speech model adaptation mechanism for adapting an existing speech model to obtain an adapted speech model for the second acoustic environment based on discrepancy between the original training data and the playback data, the existing speech model being derived based on the original training data collected in the first acoustic environment.

13. The system according to claim 12 , wherein the discrepancy based speech model adaptation mechanism comprises:

a discrepancy detection mechanism for detecting discrepancy between the original training data and the playback data; and

at least one of:

an acoustic model update mechanism for adapting an existing acoustic model to derive an adapted acoustic model, the existing acoustic model being trained based on the original training data, and the adapted acoustic model being part of the adapted speech models;

a background model update mechanism for updating an existing background model to derive an adapted background model, the existing background model being trained based on the original training data, and the adapted background model being part of the adapted speech models; or

a transformation update mechanism for generating a transformation based on the discrepancy and the existing speech model, the transformation being applied to transform the existing speech model to the adapted speech model.

14. A program code storage device, comprising:

a machine-readable storage medium; and

machine-readable program code, stored on the machine-readable storage medium, the machine readable program code having instructions, which when executed by a computing platform cause:

selecting at least a portion of original training data and playing back the selected data in the second acoustic environment to generate playback data, when the speech processing system is used in the second environment, the original training data being collected in the first acoustic environment and used to train the speech models;

recording the playback data in the second acoustic environment to generate recorded playback data; and

adapting the speech models from the first acoustic environment to the second acoustic environment based at least in part on the recorded playback data.

15. The device according to claim 14 , wherein adapting the speech models comprises re-training at least one of the speech models using at least the recorded playback data.

16. The device according to claim 14 , wherein the speech models comprise at least one of:

an acoustic model for describing the acoustic realization of at least one speech sound;

a background model for describing an acoustic environment; or

a transformation for altering some property of a given function when the transformation is applied to the function.

17. The device according to claim 14 , wherein adapting the speech models comprises:

estimating discrepancy between the original training data and the recorded playback data; and

adapting at least one of the speech models from the first acoustic environment to the second acoustic environment based on the discrepancy.

18. The device according to claim 14 , the instructions, when executed, further cause:

receiving input speech collected in the second acoustic environment; and

processing the input speech using the adapted speech models.

19. The device according to claim 17 , wherein adapting the speech models comprises:

generating a background model based on the original training data and the recorded playback data; and

producing a second acoustic model for the second acoustic environment based on the generated background model and a first acoustic model, the first acoustic model being generated using the original training data.

20. The device according to claim 17 , wherein adapting the speech models comprises:

deriving a transformation based on the speech models and the discrepancy; and

applying the transformation to the speech models to produce the adapted speech models.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 5, 2002
From: ZHOU, GUOJUN
To: INTEL CORPORATION
Reel/Frame 012763/0317 →
Continuity (1)
Related Publication 20030191636A1 · Oct 9, 2003