IP Library Granted Patent US 7,747,437
Granted Patent B2
US 7,747,437 · App. 11/305,825 · Granted Jun 29, 2010

N-best list rescoring in speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,747,437
App. No.
11/305,825
Granted
Jun 29, 2010
Kind
B2
Abstract

A method of speech recognition processing is described based on an N-best list of recognition hypotheses corresponding to a spoken input. Each hypothesis on the N-best list is rescored based on its rank in the rescored N-best list. The rescoring may be based on a Statistical Language Model (SLM) or Dynamic Semantic Model (DSM). One or more rescoring categories may be associated with each recognition hypotheses to affect or bias the rescoring.

Claims (43)

1. A method of speech recognition processing comprising:

in an automatic speech recognition engine:

producing with a pattern matching recognizer an N-best list of recognition hypotheses corresponding to a spoken input; and

rescoring the hypotheses in the automatic speech recognition engine to produce a rescored N-best list output from the automatic speech recognition engine;

wherein the rescoring uses a plurality of rescoring categories based on position in the rescored N-best list such that some positions in the rescored N-best list are rescored based on a first combination of rescoring categories and other positions in the rescored N-best list are rescored based on a second combination of rescoring categories.

2. A method according to claim 1 , wherein the rescoring is based on using a Statistical Language Model (SLM).

3. A method according to claim 1 , wherein the rescoring is based on using a Dynamic Semantic Model (DSM).

4. A method according to claim 1 , wherein the rescoring includes applying a bias to each rescored hypothesis depending on its rank in the rescored N-best list.

5. A method according to claim 1 , wherein the at least one rescoring category includes a category for most recently used recognition hypotheses.

6. A method according to claim 1 , wherein the at least one rescoring category includes a category for most frequently used recognition hypotheses.

7. A method according to claim 6 , wherein the at least one rescoring category includes a category for names within a geographic vicinity of one or more most frequently used names.

8. A method according to claim 1 , in which selected positions in the N-best list are reserved for recognition hypotheses of one or more selected rescoring categories.

9. A method according to claim 8 , wherein producing an N-best list initially considers only hypotheses in the one or more selected rescoring categories.

10. A method according to claim 1 , further comprising:

providing a first output of the rescored hypotheses in a selected number of the top positions in the rescored N-best list; and

in response to a user action, providing a second output of the remaining rescored hypotheses.

11. A method according to claim 1 , wherein the recognition hypotheses represent place names for a navigation system.

12. A method according to claim 11 , wherein the place names are city names.

13. A method according to claim 11 , wherein the place names are street names.

14. A method according to claim 1 , wherein the rescoring includes:

dividing the rescored N-best list into blocks, each block corresponding to a range of ranks in the rescored N-best list, the block boundaries varying depending on a metric corresponding to an expected recognition accuracy for the spoken input.

15. A method according to claim 14 , wherein the metric is based on a signal-to-noise ratio.

16. A speech recognition processing arrangement comprising:

an automatic speech recognition engine having:

means for producing an N-best list of recognition hypotheses corresponding to a spoken input; and

means for rescoring the hypotheses to produce a rescored N-best list output;

wherein the rescoring uses a plurality of rescoring categories based on position in the rescored N-best list such that some positions in the rescored N-best list are rescored based on a first combination of rescoring categories and other positions in the rescored N-best list are rescored based on a second combination of rescoring categories.

17. A speech processing arrangement according to claim 16 , wherein the means for rescoring is based on using Statistical Language Model (SLM) means.

18. A speech processing arrangement according to claim 16 , wherein the means for rescoring is based on using Dynamic Semantic Model (DSM) means.

19. A speech processing arrangement according to claim 16 , wherein the means for rescoring includes applying a bias to each rescored hypothesis depending on its rank in the rescored N-best list.

20. A speech processing arrangement according to claim 16 , wherein the at least one rescoring category includes a category for most recently used recognition hypotheses.

21. A speech processing arrangement according to claim 16 , wherein the at least one rescoring category includes a category for most frequently used recognition hypotheses.

22. A speech processing arrangement according to claim 21 , wherein the at least one rescoring category includes a category for names within a geographic vicinity of one or more most frequently used names.

23. A speech processing arrangement according to claim 16 , in which selected positions in the N-best list are reserved for recognition hypotheses of one or more selected rescoring categories.

24. A speech processing arrangement according to claim 23 , wherein the means for producing an N-best list initially considers only hypotheses in the one or more selected rescoring categories.

25. A speech processing arrangement according to claim 16 , further comprising:

means for providing a first output of the rescored hypotheses in a selected number of the top positions in the rescored N-best list; and

means for providing, in response to a user action, a second output of the remaining rescored hypotheses.

26. A speech processing arrangement according to claim 16 , wherein the recognition hypotheses represent place names for a navigation system.

27. A speech processing arrangement according to claim 26 , wherein the place names are city names.

28. A speech processing arrangement according to claim 26 , wherein the place names are street names.

29. A speech processing arrangement according to claim 16 , wherein the means for rescoring includes means for dividing the rescored N-best list into blocks, each block corresponding to a range of ranks in the rescored N-best list, the block boundaries varying depending on a metric corresponding to an expected recognition accuracy for the spoken input.

30. A speech processing arrangement according to claim 29 , wherein the metric is based on a signal-to-noise ratio.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065552/0934 →
PATENT RELEASE (REEL:017435/FRAME:0199) Recorded May 20, 2016
From: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
To: NUANCE COMMUNICATIONS, INC., AS GRANTOR; ART ADVANCED RECOGNITION TECHNOLOGIES, INC., A DELAWARE CORPORATION, AS GRANTOR; SPEECHWORKS INTERNATIONAL, INC., A DELAWARE CORPORATION, AS GRANTOR; TELELOGUE, INC., A DELAWARE CORPORATION, AS GRANTOR; DSP, INC., D/B/A DIAMOND EQUIPMENT, A MAINE CORPORATON, AS GRANTOR; SCANSOFT, INC., A DELAWARE CORPORATION, AS GRANTOR; DICTAPHONE CORPORATION, A DELAWARE CORPORATION, AS GRANTOR
Reel/Frame 038770/0824 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2006
From: VERHASSELT, JAN; DERCKS, HELMUT
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 017541/0240 →