IP Library Granted Patent US 8,843,370
Granted Patent B2
US 8,843,370 · App. 11/945,048 · Granted Sep 23, 2014

Joint discriminative training of multiple speech recognizers

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,843,370
App. No.
11/945,048
Granted
Sep 23, 2014
Kind
B2
Abstract

Adjusting model parameters is described for a speech recognition system that combines recognition outputs from multiple speech recognition processes. Discriminative adjustments are made to model parameters of at least one acoustic model based on a joint discriminative criterion over multiple complementary acoustic models to lower recognition word error rate in the system.

Claims (17)

1. A method of adjusting model parameters in a speech recognition system comprising:

in a speech recognition system executed by computer system that combines recognition outputs from a plurality of parallel speech recognition processes that operate on a given speech input word sequence and produce different competing recognition outputs which are then combined to determine a final recognition output;

wherein the speech recognition processes are complementary so as to produce different recognition errors on a given same speech input;

performing a discriminative adjustment process including:

i. selecting at least one acoustic model of the system, and

ii. adjusting a plurality of model parameters of the selected acoustic model based on a joint discriminative criterion over a plurality of complementary acoustic models to lower combined recognition WER in the system over only independently discriminatively trained recognition systems.

2. A method according to claim 1 , wherein the complementary acoustic models include the at least one acoustic model.

3. A method according to claim 1 , wherein the speech recognition system combines recognition outputs based on a Confusion Network Combination (CNC) approach.

4. A method according to claim 1 , wherein the speech recognition system combines recognition outputs based on a Recognizer Output Voting for Error Reduction (ROVER) approach.

5. A method according to claim 1 , wherein the joint criterion is based on Minimum Classification Error (MCE) training.

6. A method according to claim 1 , wherein the joint criterion is based on Maximum Mutual Information Estimation (MMIE) training.

7. A method according to claim 1 , wherein adjusting a plurality of model parameters includes using a Gradient Descent algorithm.

8. A method according to claim 1 , wherein adjusting a plurality of model parameters includes using an Extended Baum-Welch algorithm.

9. A method according to claim 1 , wherein the lower recognition word error rate reaches at least a 5% word error rate reduction.

10. A method according to claim 1 , wherein discriminative adjusting a plurality of model parameters includes performing parameter adaptation using at least one linear transformation.

11. A method according to claim 1 , wherein discriminative adjusting a plurality of model parameters includes performing parameter adaptation using at least one non-linear transformation.

12. A speech recognition system containing acoustic models having model parameters adjusted by the method according to any of claims 1 - 11 .

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065566/0013 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2007
From: WILLETT, DANIEL; HE, CHUANG
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 020275/0865 →