IP Library Granted Patent US 8,095,365
Granted Patent B2
US 8,095,365 · App. 12/328,436 · Granted Jan 10, 2012

System and method for increasing recognition rates of in-vocabulary words by improving pronunciation modeling

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,095,365
App. No.
12/328,436
Granted
Jan 10, 2012
Kind
B2
Abstract

The present disclosure relates to systems, methods, and computer-readable media for generating a lexicon for use with speech recognition. The method includes receiving symbolic input as labeled speech data, overgenerating potential pronunciations based on the symbolic input, identifying potential pronunciations in a speech recognition context, and storing the identified potential pronunciations in a lexicon. Overgenerating potential pronunciations can include establishing a set of conversion rules for short sequences of letters, converting portions of the symbolic input into a number of possible lexical pronunciation variants based on the set of conversion rules, modeling the possible lexical pronunciation variants in one of a weighted network and a list of phoneme lists, and iteratively retraining the set of conversion rules based on improved pronunciations. Symbolic input can include multiple examples of a same spoken word. Speech data can be labeled explicitly or implicitly and can include words as text and recorded audio.

Claims (42)

1. A method comprising:

receiving, by a processor, symbolic input as labeled speech data;

overgenerating potential pronunciations based on the symbolic input, wherein the overgenerating comprises:

establishing a set of conversion rules; and

converting portions of the symbolic input into a number of possible lexical pronunciation variants based on the set of conversion rules;

identifying potential pronunciations in a speech recognition context to yield identified potential pronunciations; and

storing the identified potential pronunciations in a lexicon.

2. The method of claim 1 , wherein symbolic input includes multiple examples of a same spoken word.

3. The method of claim 1 , wherein speech data is labeled explicitly.

4. The method of claim 1 , wherein speech data is labeled implicitly.

5. The method of claim 1 , wherein labeled speech data includes words as text and recorded audio.

6. The method of claim 1 , wherein overgenerating potential pronunciations further comprises:

modeling the possible lexical pronunciation variants in one of a weighted network and a list of phoneme lists.

7. The method of claim 1 , the method further comprising iteratively retraining the set of conversion rules based on improved pronunciations.

8. The method of claim 1 , wherein identifying potential pronunciations is based on a threshold.

9. A system for generating a lexicon for use with speech recognition, the system comprising:

a processor;

a first module configured to cause the processor to receive symbolic input as labeled speech data;

a second module configured to cause the processor to overgenerate potential pronunciations based on the symbolic input, wherein the second module further causes the processor to:

establish a set of conversion rules; and

convert portions of the symbolic input into a number of possible lexical pronunciation variants based on the set of conversion rules;

a third module configured to cause the processor to identify potential pronunciations in a speech recognition context to yield identified potential pronunciations; and

a fourth module configured to cause the processor to store the identified potential pronunciations in a lexicon.

10. The system of claim 9 , wherein symbolic input includes multiple examples of a same spoken word.

11. The system of claim 9 , wherein speech data is labeled explicitly.

12. The system of claim 9 , wherein speech data is labeled implicitly.

13. The system of claim 9 , wherein labeled speech data includes words as text and recorded audio.

14. The system of claim 9 , wherein the second module is further configured to control the processor to

model the possible lexical pronunciation variants in one of a weighted network and a list of phoneme lists.

15. The system of claim 9 , the system further comprising a fifth module configured to control the processor to iteratively retrain the set of conversion rules based on improved pronunciations.

16. A non-transitory computer-readable medium storing a computer program having instructions for generating a lexicon for use with speech recognition, the instructions comprising:

receiving symbolic input as labeled speech data;

overgenerating potential pronunciations based on the symbolic input, wherein the overgenerating comprises:

establishing a set of conversion rules; and

converting portions of the symbolic input into a number of possible lexical pronunciation variants based on the set of conversion rules;

identifying potential pronunciations in a speech recognition context to yield identified potential pronunciations; and

storing the identified potential pronunciations in a lexicon.

17. The non-transitory computer-readable medium of claim 16 , wherein symbolic input includes multiple examples of a same spoken word.

18. The non-transitory computer-readable medium of claim 16 , wherein speech data is labeled explicitly.

19. The non-transitory computer-readable medium of claim 16 , wherein speech data is labeled implicitly.

20. The non-transitory computer-readable medium of claim 16 , wherein overgenerating potential pronunciations further comprises:

modeling the possible lexical pronunciation variants in one of a weighted network and a list of phoneme lists.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2008
From: CONKIE, ALISTAIR D.; GILBERT, MAZIN; LJOLJE, ANDREJ
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 021955/0631 →