IP Library Granted Patent US 8,532,993
Granted Patent B2
US 8,532,993 · App. 13/539,996 · Granted Sep 10, 2013

Speech recognition based on pronunciation modeling

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,532,993
App. No.
13/539,996
Granted
Sep 10, 2013
Kind
B2
Abstract

A system and method for performing speech recognition is disclosed. The method comprises receiving an utterance, applying the utterance to a recognizer with a language model having pronunciation probabilities associated with unique word identifiers for words given their pronunciations and presenting a recognition result for the utterance. Recognition improvement is found by moving a pronunciation model from a dictionary to the language model.

Claims (38)

1. A method comprising:

approximating transcribed speech, via a processor, using a phonemic transcription dataset associated with a speaker, to yield a language model, where the phonemic transcription dataset is based on a pronunciation model of the speaker;

incorporating, into the language model, pronunciation probabilities associated with respective unique labels for each different pronunciation of a word, wherein the respective unique label for a most frequent word indicates a special status in the language model; and

after incorporating the pronunciation probabilities into the language model, recognizing an utterance using the language model.

2. The method of claim 1 , further comprising:

removing the pronunciation probabilities from a pronunciation dictionary.

3. The method of claim 1 , wherein the language model is generated by modeling pronunciation dependencies across word boundaries.

4. The method of claim 1 , wherein one of contextual dependencies and consistency in pronunciation style exist throughout the utterance.

5. The method of claim 1 , wherein the pronunciation probabilities in the language model further comprise pronunciation dependent word pairs as lexical items that change a behavior of the language model to approximate higher order n-gram language models.

6. The method of claim 1 , further comprising:

creating a wide context pronunciation model based on having the pronunciation probabilities in the language model; and

determining a probability of observing a particular word in the utterance using the wide context pronunciation model.

7. The method of claim 1 , wherein the pronunciation probabilities comprise a set of most frequent words each with more than one pronunciation alternative.

8. The method of claim 7 , wherein the more than one pronunciation alternative is in the language model.

9. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed on the processor, cause the processor to perform operations comprising:

approximating transcribed speech using a phonemic transcription dataset associated with a speaker, to yield a language model, where the phonemic transcription dataset is based on a pronunciation model of the speaker;

incorporating, into the language model, pronunciation probabilities associated with respective unique labels for each different pronunciation of a word, wherein the respective unique label for a most frequent word indicates a special status in the language model; and

after incorporating the pronunciation probabilities into the language model, recognizing an utterance using the language model.

10. The system of claim 9 , wherein the computer-readable storage medium has additional instructions stored which result in the operations further comprising:

removing the pronunciation probabilities from a pronunciation dictionary.

11. The system of claim 9 , wherein the language model is generated by modeling pronunciation dependencies across word boundaries.

12. The system of claim 9 , wherein one of contextual dependencies and consistency in pronunciation style exist throughout the utterance.

13. The system of claim 9 , wherein the pronunciation probabilities in the language model further comprise pronunciation dependent word pairs as lexical items that change a behavior of the language model to approximate higher order n-gram language models.

14. The system of claim 9 , the computer-readable storage medium having additional instructions stored which result in the operations further comprising:

creating a wide context pronunciation model based on having the pronunciation probabilities in the language model; and

determining a probability of observing a particular word in the utterance using the wide context pronunciation model.

15. The system of claim 9 , wherein the pronunciation probabilities comprise a set of most frequent words each with more than one pronunciation alternative.

16. The system of claim 15 , wherein the more than one pronunciation alternative is in the language model.

17. A computer-readable storage device having instructions stored which, when executed on a processor, cause the processor to perform operations comprising:

approximating transcribed speech using a phonemic transcription dataset associated with a speaker, to yield a language model, where the phonemic transcription dataset is based on a pronunciation model of the speaker;

incorporating, into the language model, pronunciation probabilities associated with respective unique labels for each different pronunciation of a word, wherein the respective unique label for a most frequent word indicates a special status in the language model; and

after incorporating the pronunciation probabilities into the language model, recognizing an utterance using the language model.

18. The computer-readable storage device of claim 17 , the computer-readable storage device having additional instructions stored which result in the operations further comprising:

removing the pronunciation probabilities from a pronunciation dictionary.

19. The computer-readable storage device of claim 17 , wherein the language model is generated by modeling pronunciation dependencies across word boundaries.

20. The computer-readable storage device of claim 17 , wherein one of contextual dependencies and consistency in pronunciation style exist throughout the utterance.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065532/0152 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2015
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 036737/0479 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2015
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 036737/0686 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 3, 2012
From: LJOLJE, ANDREJ
To: AT&T CORP.
Reel/Frame 028482/0311 →