IP Library Granted Patent US 8,417,528
Granted Patent B2
US 8,417,528 · App. 13/366,096 · Granted Apr 9, 2013

Speech recognition system with huge vocabulary

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,417,528
App. No.
13/366,096
Granted
Apr 9, 2013
Kind
B2
Abstract

The invention deals with speech recognition, such as a system for recognizing words in continuous speech. A speech recognition system is disclosed which is capable of recognizing a huge number of words, and in principle even an unlimited number of words. The speech recognition system comprises a word recognizer for deriving a best path through a word graph, and wherein words are assigned to the speech based on the best path. The word score being obtained from applying a phonemic language model to each word of the word graph. Moreover, the invention deals with an apparatus and a method for identifying words from a sound block and to computer readable code for implementing the method.

Claims (28)

1. A speech recognition system for identifying words from a sound block, the speech recognition system comprising:

a speech recognizer including one or more processors configured to derive a best path through a word graph comprising a plurality of words, wherein each of the plurality of words has an assigned word score and phonetic transcription, and wherein words in the word graph are assigned to the sound block based on the best path, wherein the word score of at least one word in the word graph was obtained by applying a phonemic language model to at least one phoneme of the at least one word of the word graph, wherein the phonemic language model includes information about at least one allowable sequence of phonemes within a word.

2. The speech recognition system according to claim 1 , wherein the speech recognition system is based on a lexicon of allowed words comprising more than 200,000 words.

3. The speech recognition system according to claim 1 , wherein the one or more processors are further configured to extract from the sound block a phoneme graph, the phoneme graph assigning a phoneme to each edge, and wherein the phonetic transcription of the words in the word graph is based on the phoneme graph.

4. The speech recognition system according to claim 1 , wherein an acoustic phoneme score is assigned to each phoneme.

5. The speech recognition system according to claim 3 , wherein the one or more processors are further configured to convert the phoneme graph to a word-phoneme graph, the word-phoneme graph assigning a word and associated phonetic transcription to each edge.

6. The speech recognition system according to claim 5 , wherein phoneme sequence hypotheses are determined and added to the phoneme graph thereby providing an extended phoneme graph, and wherein the word-phoneme graph is based on the extended phoneme graph.

7. The speech recognition system according to claim 5 , wherein the extended phoneme graph is filtered by applying a lexicon of allowed words, so as to remove phoneme sequences of the extended phoneme graph comprising words which are not present in the lexicon.

8. The speech recognition system according to claim 5 , wherein a time-synchronous word-phoneme graph is provided, and wherein words having no connection either forward or backward in time are removed from the word-phoneme graph.

9. The speech recognition system according to claim 5 , wherein the one or more processors are further configured to convert the word-phoneme graph to the word graph, the word graph assigning a word to each edge.

10. The speech recognition system according to claim 1 , wherein the phonemic language model is an m-gram language model or a compact variagram.

11. A method of identifying words from a sound block, the method comprising:

deriving, with a speech recognizer including at least one processor, a best path through a word graph comprising a plurality of words, wherein each of the plurality of words has an assigned word score and phonetic transcription, and wherein words in the word graph are assigned to the sound block based on the best path, wherein the word score of at least one word in the word graph was obtained by applying a phonemic language model to at least one phoneme of the at least one word of the word graph, wherein the phonemic language model includes information about at least one allowable sequence of phonemes within a word.

12. The method of claim 11 , further comprising:

extracting from the sound block, a phoneme graph, the phoneme graph assigning a phoneme to each edge, and wherein the phonetic transcriptions of the words in the word graph are based on the phoneme graph.

13. The method of claim 12 , further comprising:

converting the phoneme graph to a word-phoneme graph, the word-phoneme graph assigning a word and associated phonetic transcription to each edge.

14. The method according to claim 11 , further comprising:

assigning an acoustic phoneme score to each phoneme.

15. The method according to claim 13 , further comprising:

determining and adding phoneme sequence hypotheses to the phoneme graph thereby providing an extended phoneme graph, and wherein the word-phoneme graph is based on the extended phoneme graph.

16. The method according to claim 13 , further comprising:

filtering the extended phoneme graph by applying a lexicon of allowed words, so as to remove phoneme sequences of the extended phoneme graph comprising words which are not present in the lexicon.

17. The method according to claim 13 , wherein the word-phoneme graph includes time-synchronous information, the method further comprising:

removing from the word-phoneme graph based, at least in part, on the time-synchronous information, words having no connection either forward or backward in time.

18. The method according to claim 13 , further comprising:

converting the word-phoneme graph to the word graph, the word graph assigning a word to each edge.

19. The method according to claim 11 , wherein the phonemic language model is an m-gram language model or a compact variagram.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065533/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 17, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065258/0453 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2013
From: NUANCE COMMUNICATIONS AUSTRIA GMBH
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 030409/0431 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2012
From: KONINKLIJKE PHILIPS ELECTRONICS N.V.
To: NUANCE COMMUNICATIONS AUSTRIA GMBH
Reel/Frame 028877/0752 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2012
From: SAFFER, ZSOLT
To: KONINKLIJKE PHILIPS ELECTRONICS N.V.
Reel/Frame 028885/0308 →