IP Library Granted Patent US 8,301,445
Granted Patent B2
US 8,301,445 · App. 12/625,819 · Granted Oct 30, 2012

Speech recognition based on a multilingual acoustic model

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,301,445
App. No.
12/625,819
Granted
Oct 30, 2012
Kind
B2
Abstract

Embodiments of the invention relate to methods for generating a multilingual acoustic model. A main acoustic model comprising a main acoustic model having probability distribution functions and a probabilistic state sequence model including first states is provided to a processor. At least one second acoustic model including probability distribution functions and a probabilistic state sequence model including states is also provided to the processor. The processor replaces each of the probability distribution functions of the at least one second acoustic model by one of the probability distribution functions and/or each of the states of the probabilistic state sequence model of the at least one second acoustic model with the state of the probabilistic state sequence model of the main acoustic model based on a criteria set to obtain at least one modified second acoustic model. The criteria set may be a distance measurement. The processor then combines the main acoustic model and the at least one modified second acoustic model to obtain the multilingual acoustic model.

Claims (52)

1. A computer-implemented method for generating a multilingual acoustic model for use in a speech recognition system, comprising:

providing to a processor from memory a main acoustic model for a first language including a set of probability distribution functions and a probabilistic state sequence model;

providing to the processor from memory at least one second acoustic model for a second language including a set of probability distribution functions and a probabilistic state sequence model;

in a computer process, replacing each of the probability distribution functions of the at least second acoustic model by one of the probability distribution functions from the main acoustic model and/or each state of the probabilistic state sequence model from the second acoustic model by a state of the probabilistic state sequence model from the main acoustic model based upon a criteria set to form a modified second acoustic model;

in a computer process, replacing each of the second probability distribution functions of the at least one second acoustic model by the respective closest one of the first probability distribution functions to obtain a first modified second acoustic model;

in a computer process, replacing each of the second states of the second probabilistic state sequence model of the at least one second acoustic model with the respective closest state of the first probabilistic state sequence model of the main acoustic model to obtain a second modified second acoustic model;

in a computer process, weighting the first modified second acoustic model by a first weight;

in a computer process, weighting the second modified second acoustic model by a second weight; and

in a computer process, combining the first modified second acoustic model weighted by the first weight and the second modified second acoustic model weighted by the second weight and the main acoustic model to obtain the multilingual acoustic model.

2. A computer-implemented method for generating a multilingual acoustic model for use in a speech recognition system, comprising:

providing to a processor from memory a main acoustic model for a first language including a set of probability distribution functions and a probabilistic state sequence model;

providing to the processor from memory at least one second acoustic model for a second language including a set of probability distribution functions and a probabilistic state sequence model;

in a computer process, replacing the second probabilistic state sequence model of the at least one second acoustic model by the closest probabilistic state sequence model of the main acoustic model to obtain a first modified second acoustic model;

in a computer process, replacing each of the second probability distribution functions of the at least one second acoustic model by the respective closest one of the first probability distribution functions or replacing each of the second states of the second probabilistic state sequence model of the at least one second acoustic model with the respective closest state of the first probabilistic state sequence model of the main acoustic model to obtain a second modified second acoustic model;

in a computer process, weighting the first modified second acoustic model by a first weight;

in a computer process, weighting the second modified second acoustic model by a second weight; and

in a computer process, combining the first modified second acoustic model weighted by the first weight and the second modified second acoustic model weighted by the second weight and the main acoustic model to obtain the multilingual acoustic model.

3. The computer implemented method according to claim 2 , wherein the first weight is chosen between 0.4 and 0.6 and the second weight is between 0.4 and 0.6.

4. A computer-implemented method for generating a speech recognizer comprising a multilingual acoustic model, comprising:

providing to a processor from memory a main acoustic model for a first language including a set of probability distribution functions and a probabilistic state sequence model;

providing to the processor from memory at least one second acoustic model for a second language including a set of probability distribution functions and a probabilistic state sequence model;

in a computer process, replacing the second probabilistic state sequence model of the at least one second acoustic model by the closest probabilistic state sequence model of the main acoustic model to obtain a first modified second acoustic model;

in a computer process, replacing each of the second probability distribution functions of the at least one second acoustic model by the respective closest one of the first probability distribution functions to obtain a second modified second acoustic model;

in a computer process, replacing each of the second states of the second probabilistic state sequence model of the at least one second acoustic model with the respective closest state of the first probabilistic state sequence model of the main acoustic model to obtain a third modified second acoustic model;

in a computer process, weighting the first modified second acoustic model by a first weight;

in a computer process, weighting the second modified second acoustic model by a second weight;

in a computer process, weighting the third modified second acoustic model by a third weight; and

in a computer process, combining the first modified second acoustic model weighted by the first weight, the second modified second acoustic model weighted by the second weight, the third modified second acoustic model weighted by the third weight and the main acoustic model to obtain the multilingual acoustic model.

5. A computer-implemented method for generating a multilingual acoustic model for use in a speech recognition system, comprising:

providing to a processor from memory a main acoustic model for a first language including a set of probability distribution functions and a probabilistic state sequence model;

providing to the processor from memory at least one second acoustic model for a second language including a set of probability distribution functions and a probabilistic state sequence model;

in a computer process, replacing each of the probability distribution functions of the at least second acoustic model by one of the probability distribution functions from the main acoustic model and/or each state of the probabilistic state sequence model from the second acoustic model by a state of the probabilistic state sequence model from the main acoustic model based upon a criteria set to form a modified second acoustic model; and

in a computer process, combining the main acoustic model and the at least one modified second acoustic model to form the multilingual acoustic model;

wherein the criteria set is a distance measurement determined based on the Mahalanobis distance between first and second probability distribution functions.

6. A computer program product comprising a non-transitory computer readable medium having executable computer code thereon for generating a speech recognizer comprising a multilingual acoustic model, the computer code comprising:

computer code for retrieving a main acoustic model for a first language including first probability distribution functions and a probabilistic state sequence model having states;

computer code for retrieving at least one second acoustic model for a second language including probability distribution functions and a probabilistic state sequence model including states;

computer code for replacing each of the probability distribution functions of the at least one second acoustic model by the respective closest one of the first probability distribution functions to obtain a first modified second acoustic model;

computer code for replacing each of the states of the probabilistic state sequence model of the at least one second acoustic model with the respective closest state of the probabilistic state sequence model of the main acoustic model to obtain a second modified second acoustic model;

computer code for weighting the first modified second acoustic model by a first weight;

computer code for weighting the second modified second acoustic model by a second weight; and

computer code for combining the first modified second acoustic model weighted by the first weight and the second modified second acoustic model weighted by the second weight and the main acoustic model to obtain the multilingual acoustic model.

7. A computer program product comprising a non-transitory computer readable medium having executable computer code thereon for generating a speech recognizer comprising a multilingual acoustic model, the computer code comprising:

computer code for retrieving a main acoustic model for a first language including first probability distribution functions and a probabilistic state sequence model having states;

computer code for retrieving at least one second acoustic model for a second language including probability distribution functions and a probabilistic state sequence model including states;

computer code for replacing the probabilistic state sequence model of the at least one second acoustic model by the closest probabilistic state sequence model of the main acoustic model to obtain a first modified second acoustic model;

computer code for replacing each of the probability distribution functions of the at least one second acoustic model by the respective closest one of the probability distribution functions from the main acoustic model to obtain a second modified second acoustic model;

computer code for replacing each of the states of the probabilistic state sequence model of the at least one second acoustic model with the respective closest state of the probabilistic state sequence model of the main acoustic model to obtain a third modified second acoustic model;

computer code for weighting the first modified second acoustic model by a first weight;

computer code for weighting the second modified second acoustic model by a second weight;

computer code for weighting the third modified second acoustic model by a third weight; and

computer code for combining the first modified second acoustic model weighted by the first weight, the second modified second acoustic model weighted by the second weight, the third modified second acoustic model weighted by the third weight and the main acoustic model to obtain the multilingual acoustic model.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2010
From: GRUHN, RAINER; RAAB, MARTIN; BRUECKNER, RAYMOND
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 024058/0779 →