IP Library Granted Patent US 10,360,904
Granted Patent B2
US 10,360,904 · App. 15/309,282 · Granted Jul 23, 2019

Methods and apparatus for speech recognition using a garbage model

Inventors: Cosmin Popovici (Giaveno, IT); Kenneth W. D. Smith (London, GB); Petrus C. Cools (Kerkrade, NL)
Assignee: Nuance Communications, Inc.
G10L15/187G10L15/02G10L15/32G10L2015/025G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,360,904
App. No.
15/309,282
Granted
Jul 23, 2019
Kind
B2
Abstract

Methods and apparatus for performing speech recognition using a garbage model. The method comprises receiving audio comprising speech and processing at least some of the speech using a garbage model to produce a garbage speech recognition result. The garbage model includes a plurality of sub-words, each of which corresponds to a possible combination of phonemes in a particular language.

Claims (39)

1. A method of performing speech recognition using a garbage model grammar configured to recognize out-of-grammar words that are not included in a first grammar, the method comprising:

processing, by a speech recognition system including at least one processor, at least some of input speech using the garbage model grammar to produce a first speech recognition result including at least one out-of-grammar word that is not included in the first grammar,

wherein the garbage model grammar includes a plurality of sub-words, each of which is one or more characters in a phonetic alphabet and corresponds to a possible combination of phonemes in a particular language, and processing the at least some of the input speech using the garbage model grammar comprises comparing a phonetic transcription of the at least some of the input speech to the plurality of sub-words of the garbage model grammar to identify the at least one out-of-grammar word of the first speech recognition result,

wherein the plurality of sub-words included in the garbage model grammar include only a subset of all possible phoneme combinations in the particular language and wherein each of the plurality of sub-words in the garbage model grammar includes a vowel.

2. The method of claim 1 , further comprising:

determining whether to process the at least some of the input speech using the garbage model grammar; and

processing the at least some of the input speech using the garbage model grammar when it is determined to process the at least some of the input speech using the garbage model grammar.

3. The method of claim 2 , further comprising:

recognizing the input speech using a first speech recognition model to produce a second speech recognition result having an associated confidence value, wherein the first speech recognition model is different than the garbage model grammar; and

wherein determining whether to process the at least some of the input speech using the garbage model grammar comprises determining to process the at least some of the input speech using the garbage model grammar when the confidence value is less than a threshold value.

4. The method of claim 3 , wherein the first speech recognition model comprises an application grammar.

5. The method of claim 3 , wherein the first speech recognition model comprises a keyword spotting grammar.

6. The method of claim 2 , wherein the input speech comprises a beginning portion, a middle portion, and an end portion, and wherein determining whether to process the at least some of the input speech using the garbage model grammar comprises determining to recognize the beginning portion and/or the end portion of the input speech using the garbage model grammar.

7. The method of claim 1 , wherein the first speech recognition result is associated with a first confidence value, the method further comprising:

processing the at least some of the input speech using a first speech recognition model to produce a second speech recognition result having a second confidence value, wherein the first speech recognition model is different than the garbage model grammar; and

identifying the at least some of the input speech as out-of-grammar speech when the first confidence value is greater than the second confidence value.

8. The method of claim 1 , wherein the method further comprises:

receiving a first transcription associated with the input speech;

processing the input speech using the garbage model grammar to produce a second transcription; and

determining that the first transcription is incorrect based, at least in part, on a comparison of the first transcription and the second transcription.

9. The method of claim 1 , wherein the speech recognition system, for an 80% hit rate achieves a false alarm rate of less than 10 false alarms per hour of received audio.

10. The method of claim 1 , wherein the speech recognition system, for an 80% hit rate achieves a false alarm rate of less than 5 false alarms per hour of received audio.

11. The method of claim 1 , wherein the speech recognition system, for an 80% hit rate achieves a false alarm rate of less than 1 false alarm per hour of received audio.

12. The method of claim 1 , wherein the speech recognition system, for an 80% hit rate achieves a false alarm rate of less than 0.5 false alarms per hour of received audio.

13. The method of claim 1 , wherein at least one sub-word in the plurality of sub-words in the garbage model grammar is a word.

14. The method of claim 13 , wherein the at least one sub-word that is a word includes only one vowel or only one semi-vowel.

15. The method of claim 1 , wherein processing the at least some of the input speech using the garbage model grammar to produce the first speech recognition result comprises processing the at least some of the input speech using a lexicon including transcriptions for the plurality of sub-words in the garbage model grammar.

16. A speech recognition system, comprising:

at least one input interface configured to receive audio comprising speech;

a garbage model grammar configured to recognize out-of-grammar words that are not included in a first grammar, wherein the garbage model grammar includes a plurality of sub-words, each of which is one or more characters in a phonetic alphabet, includes a vowel, and corresponds to a possible combination of phonemes in a particular language, wherein the plurality of sub-words included in the garbage model grammar include only a subset of all possible phoneme combinations in the particular language; and

at least one processor programmed to process at least some of the speech using the garbage model grammar to produce a first speech recognition result including at least one out-of-grammar word that is not included in the first grammar, wherein the processing comprises comparing a phonetic transcription of the at least some of the speech to the plurality of sub-words of the garbage model grammar to identify the at least one out-of-grammar word of the first speech recognition result.

17. The speech recognition system of claim 16 , further comprising a first speech recognition model different than the garbage model grammar, wherein the first speech recognition result is associated with a first confidence value, and wherein the at least one processor is further programmed to:

process the at least some of the speech using the first speech recognition model to produce a second speech recognition result having a second confidence value; and

identify the at least some of the speech as out-of-grammar speech when the first confidence value is greater than the second confidence value.

18. The speech recognition system of claim 16 , wherein the at least one input interface is further configured to receive a first transcription associated with the speech; and wherein the at least one processor is further programmed to:

process the speech using the garbage model grammar to produce a second transcription; and

determine that the first transcription is incorrect based, at least in part, on a comparison of the first transcription and the second transcription.

19. A non-transitory computer-readable medium including a plurality of instructions that, when executed by at least one processor of a speech recognition system, perform a method of performing speech recognition using a garbage model grammar configured to recognize out-of-grammar words that are not included in a first grammar, the method comprising:

processing at least some of input speech using the garbage model grammar to produce a first speech recognition result including at least one out-of-grammar word that is not included in the first grammar, wherein the garbage model grammar includes a plurality of sub-words, each of which is one or more characters in a phonetic alphabet, includes a vowel, and corresponds to a possible combination of phonemes in a particular language, and wherein the plurality of sub-words included in the garbage model grammar include only a subset of all possible phoneme combinations in the particular language, and wherein the processing comprises comparing a phonetic transcription of the at least some of the input speech to the plurality of sub-words of the garbage model grammar to identify the at least one out-of-grammar word of the first speech recognition result.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065566/0013 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2017
From: POPOVICI, COSMIN; SMITH, KENNETH W.D.; COOLS, PETRUS C.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041054/0755 →
Continuity (1)
Related Publication 20170076718A1 · Mar 16, 2017