IP Library Granted Patent US 7,313,526
Granted Patent B2
US 7,313,526 · App. 10/950,092 · Granted Dec 25, 2007

Speech recognition using selectable recognition modes

Assignee: Voice Signal Technologies, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,313,526
App. No.
10/950,092
Granted
Dec 25, 2007
Kind
B2
Abstract

The present invention relates to speech recognition using selectable recognition modes. This includes innovations such as: large vocabulary speech recognition programming that supplies recognized words to external program as they are recognized, and allows a user to select between large vocabulary recognition of an utterance with and without language context from the prior utterance independently of state of the external program; allowing a user to select between continuous and discrete speech recognition that use substantially the same vocabulary; allowing a user to select between continuous and discrete large-vocabulary speech recognition modes; allowing a user to select between at least two different alphabetic entry speech recognition modes; and allowing a user to select from among four or more of the following recognitions modes when creating text: a large-vocabulary mode, an alphabetic entry mode, a number entry mode, and a punctuation entry mode.

Claims (97)

1. A computerized method of performing speech recognition comprising:

using speech recognition programming for:

providing a user interface which allows a user to select between generating a first and a second user input;

responding to the generation of the first user input by performing large vocabulary speech recognition on one or more utterances in a prior-language-context-dependent mode, which recognizes at least the first word of an utterance depending in part on a language model context created by a previously recognized word from the previous utterance, if any; and

responding to the generation of the second user input by performing large vocabulary speech recognition on one or more utterances in a prior-language-context-independent mode, which recognizes at least the first word of an utterance substantially independently of any language model context created by a previously recognized word from the previous utterance, if any;

wherein:

as words are recognized by said speech recognition programming in both of said recognition modes such words are output to programming external to said speech recognition programming for use by said external programming; and

the response by said speech recognition programming to said first and second inputs by switching recognition modes is independent of the state of said external programming.

2. A method as in claim 1 wherein:

the user interface includes a first button and a second button, where said buttons can be either hardware or software buttons;

the first user input is generated by pressing the first button; and

the second user input is generated by pressing the second button.

3. A method as in claim 1 wherein the prior-language-context-independent mode uses language context probabilities within an utterance, causing the recognition of a word in a given utterance to depend on the identity of the one or more words, if any, recognized before it in said given utterance.

4. A method as in claim 1 wherein said method is performed by a software input panel in Microsoft Windows CE.

5. A computerized method of performing speech recognition comprising:

providing a user interface which allows a user to select between generating a first and a second user input;

responding to the generation of the first user input by selecting a continuous speech recognition mode which performs continuous speech recognition on speech sounds using a given vocabulary;

responding to the generation of the second user input by selecting a discrete recognition mode which performs discrete recognition on speech sounds using substantially the same given vocabulary; and

responding to speech sounds by performing recognition upon them using the currently selected speech recognition mode;

wherein:

the user can switch between the use of continuous and discrete recognition by selecting one of said user inputs;

the one or more input devices include a first button and a second button;

the first user input is generated by pressing the first button;

The second user input is generated by pressing the second button;

touching the first or second button causes its respective recognition mode to start from substantially the start of the touching of such a button and to terminate by the next detection of an end of utterance;

the discrete recognition is limited to the recognition of the one or more vocabulary word candidates with the best scoring match against the utterance whose end is detected after the touching of said button; and

the continuous recognition mode is not so limited;

so that said discrete recognition mode is limited to outputting only one single vocabulary word as the best scoring recognition candidate for the recognition of the given utterance and the continuous recognition mode can output a sequence of multiple words for the recognition of the given utterance.

6. A method as in claim 5 wherein the given vocabulary is a large vocabulary.

7. A method as in claim 5 wherein the given vocabulary is an alphabetic input vocabulary.

8. A method as in claim 5 wherein:

said user interface allows a user to select between generating a third and a fourth input independently from the selection of the first and second input; and

said method further includes responding to said third and fourth inputs, respectively, by selecting as said given vocabulary a first vocabulary or a second vocabulary;

whereby the user can separately switch between recognition vocabularies and between THE discrete and continuous recognition modes.

9. A method as in claim 8 wherein said first and second vocabulary are a large vocabulary of words and an alphabetic input vocabulary, respectively.

10. A method as in claim 8 wherein said first and second vocabulary are two different alphabetic input vocabularies.

11. A method as in claim 5 wherein:

the user interface provided includes a first button and a second button;

the first user input is generated by pressing the first button; and

the second user input is generated by pressing the second button.

12. A method as in claim 11 wherein:

touching the first and second buttons causes their respective recognition mode to recognize from substantially the start of the touching of such a button until the next end of utterance is detected;

the discrete recognition is limited to the recognition of the one or more single vocabulary word recognition candidates with the best scoring match against the given utterance whose end is detected after the touching of said button; and

the continuous recognition mode is not limited;

so that said discrete recognition mode is limited to outputting only one single vocabulary word as the best scoring recognition candidate for the recognition of the given utterance and the continuous recognition mode can output a sequence of multiple words for the recognition of the given utterance.

13. A method as in claim 5 wherein acoustic models used to represent words in the discrete recognition mode are different than the acoustic models used to represent the same words in the continuous recognition mode.

14. A computerized method of performing speech recognition comprising:

providing a user interface which allows a user to select between generating a first and a second user input;

responding to the generation of the first user input by switching to a first recognition mode that recognizes one or more utterances as one or more words in a first alphabetic entry vocabulary; and

responding to the generation of the second user input by switching to a second recognition mode that recognizes one or more utterances as one or more words in a second, different, alphabetic entry vocabulary;

wherein the first and second alphabetic entry vocabularies associated different letter-identifying words with individual letters of the alphabet.

15. A method as in claim 14 wherein:

the first alphabetic entry vocabulary includes the names of each letter of the alphabet and the second alphabetic entry vocabulary does not; and

the second alphabetic entry vocabulary includes one or more words that start with each letter of the alphabet and the first alphabetic entry vocabulary does not.

16. A method as in claim 14 wherein said user interface provides a separate button for generating said first and second inputs.

17. A method as in claim 16 wherein touching of each of said buttons turns on recognition in the button's associated alphabetic entry mode.

18. A method as in claim 14 wherein

said user interface enables:

a user to select a filtering mode in which word choices for the recognition of a given word are limited to word's whose spelling matches a sequence of one or more characters input by the user;

a user to enter said one or more filtering characters by voice recognition using either said first or second alphabetic entry modes; and

said first and second inputs select between whether such recognition of filtering characters is performed using said first or second alphabetic entry modes, respectively.

19. A computerized method of performing speech recognition comprising:

providing a user interface which allows a user to select between generating a first, a second, a third, or a fourth user input;

responding to the generation of the first user input by switching to performing speech recognition using a first, general purpose large vocabulary; and

responding to the generation of the second user input by switching to performing speech recognition using a second, alphabetic entry vocabulary;

responding to the generation of the third user input by switching to performing speech recognition using a third, numerical entry vocabulary;

responding to the generation of the fourth user input by switching to performing speech recognition using a fourth, punctuation entry, vocabulary; and

sequentially receiving output in the form of words produced by speech recognizing using different user selected ones of said four vocabularies and placing that output into a common text.

20. A method as in claim 19 wherein the output of speech recognition using the third vocabulary is in the form of numerical digits.

21. A method as in claim 19 wherein the output of speech recognition using the fourth vocabulary is in the form of punctuation marks.

22. A method as in claim 19 wherein the user interface provides a different button for the selection of each of the first, second, third and fourth inputs.

23. A method as in claim 22 wherein pressing the button associated with one of said four vocabularies turns on recognition using that vocabulary.

24. A computerized method of performing speech recognition comprising:

providing a user interface which allows a user to select between generating a first and a second user input;

responding to the generation of the first user input by selecting a continuous speech recognition mode which performs continuous speech recognition on speech sounds using a given vocabulary;

responding to the generation of the second user input by selecting a discrete recognition mode which performs discrete recognition on speech sounds using substantially the same given vocabulary; and

responding to speech sounds by performing recognition upon them using the currently selected speech recognition mode;

wherein:

the user can switch between the use of continuous and discrete recognition by selecting one of said user inputs;

said user interface allows a user to select between generating a third and a fourth input independently from the selection of the first and second input;

said method further includes responding to said third and fourth inputs, respectively, by selecting as said given vocabulary a first vocabulary or a second vocabulary;

whereby the user can separately switch between recognition vocabularies and between the discrete and continuous recognition modes; and

said first and second vocabulary are two different alphabetic input vocabularies.

25. A computerized method of performing speech recognition comprising:

providing a user interface which allows a user to select between generating a first and a second user input;

responding to the generation of the first user input by selecting a continuous speech recognition mode which performs continuous speech recognition on speech sounds using a given vocabulary;

responding to the generation of the second user input by selecting a discrete recognition mode which performs discrete recognition on speech sounds using substantially the same given vocabulary; and

responding to speech sounds by performing recognition upon them using the currently selected speech recognition mode;

wherein:

the user can switch between the use of continuous and discrete recognition by selecting one of said user inputs

the user interface provided includes a first button and a second button;

the first user input is generated by pressing the first button;

the second user input is generated by pressing the second button;

touching the first and second buttons causes their respective recognition mode to recognize from substantially the start of the touching of such a button until the next end of utterance is detected;

the discrete recognition is limited to the recognition of the one or more single vocabulary word recognition candidates with the best scoring match against the given utterance whose end is detected after the touching of said button; and

the continuous recognition mode is not limited;

so that said discrete recognition mode is limited to outputting only one single vocabulary word as the best scoring recognition candidate for the recognition of the given utterance and the continuous recognition mode can output a sequence of multiple words for the recognition of the given utterance.

Assignments (9)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
MERGER Recorded Sep 13, 2012
From: VOICE SIGNAL TECHNOLOGIES, INC.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 028952/0277 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2004
From: ROTH, DANIEL L.; COHEN, JORDAN R.; JOHNSTON, DAVID F.; GRABHERR, MANFRED G.
To: VOICE SIGNAL TECHNOLOGIES, INC.
Reel/Frame 015842/0644 →
Continuity (16)
Continuation In Part 1022765300 · Sep 6, 2002
Continuation In Part 1030205300 · Sep 5, 2002
Provisional Application 6031733300 · Sep 5, 2001
Provisional Application 6031743300 · Sep 5, 2001
Provisional Application 6031743100 · Sep 5, 2001
Provisional Application 6031732900 · Sep 5, 2001
Provisional Application 6031733000 · Sep 5, 2001
Provisional Application 6031733100 · Sep 5, 2001
Provisional Application 6031742300 · Sep 5, 2001
Provisional Application 6031742200 · Sep 5, 2001
Provisional Application 6031742100 · Sep 5, 2001
Provisional Application 6031743000 · Sep 5, 2001
Provisional Application 6031743200 · Sep 5, 2001
Provisional Application 6031743500 · Sep 5, 2001
Provisional Application 6031743400 · Sep 5, 2001
Related Publication 20050049880A1 · Mar 3, 2005