IP Library Granted Patent US 7,533,020
Granted Patent B2
US 7,533,020 · App. 11/063,357 · Granted May 12, 2009

Method and apparatus for performing relational speech recognition

Assignee: Nuance Communications, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,533,020
App. No.
11/063,357
Granted
May 12, 2009
Kind
B2
Abstract

A method and apparatus are provided for performing speech recognition using observable and meaningful relationships between words within a single utterance and using a structured data source as a source of constraints on the recognition process. Results from a first constrained speech recognition pass can be combined with information about the observable and meaningful word relationships to constrain or simplify subsequent recognition passes. This iterative process greatly reduces the search space required for each recognition pass, making the speech recognition process more efficient, faster and accurate.

Claims (43)

1. A method for performing speech recognition, the method comprising:

acquiring a speech signal from a user;

performing a first recognition pass by applying a first language model to said speech signal, said first language model being constrained in accordance with a structured data source; and

generating a subsequent language model based at least in part on results from said first recognition pass.

2. The method of claim 1 , further comprising:

performing a subsequent recognition pass by applying said subsequent language model; and

recognizing said speech signal.

3. The method of claim 1 , wherein said structured data source is at least one of a geographic location report, a map, a blueprint, a telephone directory, a service directory, the Internet, a mobile address book and one or more user preferences.

4. The method of claim 3 , wherein said one or more user preferences are learned.

5. The method of claim 3 , wherein said one or more user preferences are programmed by said user.

6. The method of claim 1 , wherein said first language model is constrained to a limited geographic domain encompassing said user's current geographic location.

7. The method of claim 6 , wherein said user's current geographic location is acquired prior to acquiring said speech signal.

8. The method of claim 7 , wherein said user's current geographic location is acquired periodically based on said user's travel.

9. The method of claim 1 , wherein said first language model is constrained to a geographic area specified by said user.

10. The method of claim 9 , wherein said first language model is constrained to a union of geographic areas surrounding a plurality of waypoints identified by said user.

11. A computer readable storage medium containing an executable program for performing speech recognition, where the program performs the steps of:

acquiring a speech signal from a user;

performing a first recognition pass by applying a first language model to said speech signal, said first language model being constrained in accordance with a structured data source; and

generating a subsequent language model based at least in part on results from said first recognition pass.

12. The computer readable storage medium of claim 11 , further comprising:

performing a subsequent recognition pass by applying said subsequent language model; and

recognizing said speech signal.

13. The computer readable storage medium of claim 11 , wherein said structured data source is at least one of a geographic location report, a map, a blueprint, a telephone directory, a service directory, the Internet, a mobile address book and one or more user preferences.

14. The computer readable storage medium of claim 13 , wherein said one or more user preferences are learned.

15. The computer readable storage medium of claim 13 , wherein said one or more user preferences are programmed by said user.

16. The computer readable storage medium of claim 11 , wherein said first language model is constrained to a limited geographic domain encompassing said user's current geographic location.

17. The computer readable storage medium of claim 16 , wherein said user's current geographic location is acquired prior to acquiring said speech signal.

18. The computer readable storage medium of claim 17 , wherein said user 's current geographic location is acquired periodically based on said user's travel.

19. The computer readable storage medium of claim 11 , wherein said first language model is constrained to a geographic area specified by said user.

20. The computer readable storage medium of claim 19 , wherein said first language model is constrained to a union of geographic areas surrounding a plurality of waypoints identified by said user.

21. Apparatus for performing speech recognition, said apparatus comprising:

means for acquiring a speech signal from a user;

means for performing a first recognition pass by applying a first language model to said speech signal, said first language model being constrained in accordance with a structured data source; and

means for generating a subsequent language model based at least in part on results from said first recognition pass.

22. An apparatus for performing speech recognition, the apparatus comprising:

at least one computer readable storage medium encoded with a plurality of instructions; and

at least one processer programmed by at least some of the plurality of instructions to;

acquire a speech signal from a user;

perform a first recognition pass by applying a first language model to the speech signal, the first language model being constrained in accordance with a structured data source; and

generate a subsequent language model based at least in part on results from the first recognition pass.

23. The apparatus of claim 22 , wherein the at least one processer is further programmed to:

perform a subsequent recognition pass by applying the subsequent language model; and

recognize the speech signal.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2008
From: SRI INTERNATIONAL
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 021677/0969 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNOR NAME VENKATA RAMANA RAO GADDE PREVIOUSLY RECORDED ON REEL 018811 FRAME 0366. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNOR LAST NAME FROM RAO GADDE TO GADDE. Recorded Apr 27, 2007
From: ARNOLD, JAMES; FRANDSEN, MICHAEL W.; VENKATARAMAN, ANAND; BERCOW, DOUGLAS; MYERS, GREGORY; ISRAEL, DAVID; GADDE, VENKATA RAMANA RAO; FRANCO, HORACIO; BRATT, HARRY
To: SRI INTERNATIONAL
Reel/Frame 019221/0165 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2007
From: FRANDSEN, MICHAEL W.; VENKATARAMAN, ANAND; BERCOW, DOUGLAS; MYERS, GREGORY; ISRAEL, DAVID; RAO GADDE, VENKATA RAMANA; FRANCO, HORACIO; BRATT, HARRY; ARNOLD, JAMES
To: SRI INTERNATIONAL
Reel/Frame 018811/0366 →
Continuity (2)
Continuation In Part 0996722800 · Sep 28, 2001
Related Publication 20050234723A1 · Oct 20, 2005