IP Library Granted Patent US 9,892,733
Granted Patent B2
US 9,892,733 · App. 14/283,037 · Granted Feb 13, 2018

Method and apparatus for an exemplary automatic speech recognition system

Inventor: Fathy Yassa (Soquel, CA)
Assignee: SPEECH MORPHING SYSTEMS, INC.
G10L15/26G06F17/275G10L15/00G10L15/005G10L15/183G10L15/32
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,892,733
App. No.
14/283,037
Granted
Feb 13, 2018
Kind
B2
Abstract

An exemplary computer system configured to user multiple automatic speech recognizers (ASRs) with a plurality of language and acoustic models to increase the accuracy of speech recognition.

Claims (40)

1. An automatic speech recognition system comprising:

an input interface configured to receive input speech;

a pre-processor configured to determine a language of the input speech and, based on the language of the input speech, select from among a plurality of automatic speech recognition engines configured to recognize speech of different languages, only those automatic speech recognition engines configured to recognize the language of the input speech including:

a first automatic speech recognition (ASR) engine configured to translate the input speech of the language into text according to a first language model, and output first translated text that is translated from the input speech according to the first language model and a first confidence score indicating a degree of accuracy of translating the input speech into the first translated text according to the first language model;

a second ASR engine configured to translate the input speech of the language into text according to a second language model, and output second translated text that is translated from the input speech according to the second language model and a second confidence score indicating a degree of accuracy of translating the input speech into the second translated text according to the second language model; and

a comparator configured to compare the first confidence score and the second confidence score, and output a most accurate representation of the input speech from among the first translated text and the second translated text based on a result of the comparison,

wherein the first language model is different from the second language model.

2. The automatic speech recognition system of claim 1 , wherein the first ASR engine is further configured to translate the input speech of the language into the first translated text according to the first language model and a first acoustic model,

wherein the second ASR engine is further configured to translate the input speech of the language into the second translated text according to the second language model and a second acoustic model, and

wherein the first acoustic model is different from the second acoustic model.

3. The automatic speech recognition system of claim 2 , wherein the first acoustic model is a first process of establishing a statistical representation for feature vector sequences computed from a speech waveform of the input speech, and

wherein the second acoustic model is a second process of establishing a statistical representation for feature vector sequences computed from a speech waveform of the input speech different from the first process.

4. The automatic speech recognition system of claim 3 , wherein the first process comprises first pronunciation modeling configured to describe how at least one sequence of fundamental speech units of the input speech is used to represent at least one word or phrase that is an object of speech recognition, and

wherein the second process comprises second pronunciation modeling configured to describe how at least one sequence of fundamental speech units of the input speech is used to represent at least one word or phrase that is an object of speech recognition different from the second pronunciation modeling.

5. The automatic speech recognition system of claim 1 , wherein the first ASR engine is further configured to translate the input speech of the language into the first translated text according to the first language model and a first acoustic model,

wherein the second ASR engine is further configured to translate the input speech of the language into the second translated text according to the second language model and a second acoustic model, and

wherein the first acoustic model is the same as the second acoustic model.

6. The automatic speech recognition system of claim 5 , wherein the first acoustic model and the second acoustic model are a processes of establishing a statistical representation for feature vector sequences computed from a speech waveform of the input speech.

7. The automatic speech recognition system of claim 6 , wherein the processes comprise pronunciation modeling configured to describe how at least one sequence of fundamental speech units of the input speech is used to represent at least one word or phrase that is an object of speech recognition.

8. The automatic speech recognition system of claim 1 , wherein the first language model defines a first relationship of ordering among words in phrases of the language, and

wherein the second language model defines a second relationship of ordering among words in phrases of the language.

9. An automatic speech recognition system comprising:

an input interface configured to receive input speech;

a pre-processor configured to determine a language of the input speech and, based on the language of the input speech, select from among a plurality of automatic speech recognition engines configured to recognize speech of different languages, only those automatic speech recognition engines configured to recognize the language of the input speech including:

a first automatic speech recognition (ASR) engine configured to translate the input speech of the language into text according to a first acoustic model, and output first translated text that is translated from the input speech according to the first acoustic model and a first confidence score indicating a degree of accuracy of translating the input speech into the first translated text according to the first acoustic model;

a second ASR engine configured to translate the input speech of the language into text according to a second acoustic model, and output second translated text that is translated from the input speech according to the second acoustic model and a second confidence score indicating a degree of accuracy of translating the input speech into the second translated text according to the second acoustic model; and

a comparator configured to compare the first confidence score and the second confidence score, and output a most accurate representation of the input speech from among the first translated text and the second translated text based on a result of the comparison,

wherein the first acoustic model is different from the second acoustic model.

10. The automatic speech recognition system of claim 9 , wherein the first ASR engine is further configured to translate the input speech of the language into the first translated text according to the first acoustic model and a first language model,

wherein the second ASR engine is further configured to translate the input speech of the language into the second translated text according to the second acoustic model and a second language model, and

wherein the first language model is different from the second language model.

11. The automatic speech recognition system of claim 10 , wherein the first acoustic model is a first process of establishing a statistical representation for feature vector sequences computed from a speech waveform of the input speech, and

wherein the second acoustic model is a second process of establishing a statistical representation for feature vector sequences computed from a speech waveform of the input speech different from the first process.

12. The automatic speech recognition system of claim 11 , wherein the first process comprises first pronunciation modeling configured to describe how at least one sequence of fundamental speech units of the input speech is used to represent at least one word or phrase that is an object of speech recognition, and

wherein the second process comprises second pronunciation modeling configured to describe how at least one sequence of fundamental speech units of the input speech is used to represent at least one word or phrase that is an object of speech recognition different from the second pronunciation modeling.

13. The automatic speech recognition system of claim 10 , wherein the first language model defines a first relationship of ordering among words in phrases of the language, and

wherein the second language model defines a second relationship of ordering among words in phrases of the language.

14. The automatic speech recognition system of claim 9 , wherein the first ASR engine is further configured to translate the input speech of the language into the first translated text according to the first acoustic model and a first language model,

wherein the second ASR engine is further configured to translate the input speech of the language into the second translated text according to the second acoustic model and a second language model, and

wherein the first language model is the same as the second language model.

Assignments (2)
CHANGE OF NAME Recorded Mar 29, 2016
From: SPEECH MORPHING, INC.
To: SPEECH MORPHING SYSTEMS, INC.
Reel/Frame 038123/0026 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2014
From: YASSA, FATHY
To: SPEECH MORPHING, INC.
Reel/Frame 033776/0742 →
Continuity (2)
Provisional Application 61825516 · May 20, 2013
Related Publication 20140343940A1 · Nov 20, 2014