IP Library Granted Patent US 8,332,218
Granted Patent B2
US 8,332,218 · App. 11/423,710 · Granted Dec 11, 2012

Context-based grammars for automated speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,332,218
App. No.
11/423,710
Granted
Dec 11, 2012
Kind
B2
Abstract

Methods, apparatus, and computer program products for providing a context-based grammar for automatic speech recognition, including creating by a multimodal application a context, the context comprising words associated with user activity in the multimodal application, and supplementing by the multimodal application a grammar for automatic speech recognition in dependence upon the context.

Claims (45)

1. A method of providing a context-based grammar for automatic speech recognition at a first web page, the method comprising:

creating by a multimodal application a context, the context comprising words taken from at least one second web page different from the first web page, and/or taken from at least one grammar for the at least one second web page, wherein creating the context comprises maintaining a history of at least some words from previously visited web pages, generating occurrence counts for words in the history, and selecting words for the context based at least in part on the occurrence counts; and

supplementing by the multimodal application a grammar and/or lexicon for automatic speech recognition at the first web page in dependence upon the context.

2. The method of claim 1 , wherein creating the context further comprises inserting in the context one or more words from a grammar from a previously visited multimodal web page.

3. The method of claim 1 , wherein creating the context further comprises selecting words for the context from the history.

4. The method of claim 1 , wherein creating the context further comprises generating a frequent values list for words in the history, and selecting words for the context from the history in dependence upon the frequent values list.

5. The method of claim 1 , wherein creating the context further comprises generating a list of most recently used words in the history, and selecting words for the context from the history in dependence upon the list of most recently used words.

6. The method of claim 1 , wherein creating the context further comprises generating a list of words from the history used during a predetermined period of time, and selecting words for the context from the history in dependence upon the list of words from the history used during a predetermined period of time.

7. The method of claim 1 , wherein the grammar is a VoiceXML grammar.

8. The method of claim 1 , wherein the multimodal application runs on a voice server.

9. The method of claim 1 , further comprising:

inferring at least one phoneme for at least one word in the grammar using text to speech conversion; and

adding the at least one phoneme for the at least one word to the lexicon.

10. The method of claim 1 , wherein the context comprises words taken from the at least one second web page.

11. The method of claim 1 , wherein the context comprises words taken from at least one grammar for the at least one second web page.

12. A system for providing a context-based grammar for automatic speech recognition at a first web page, the system comprising a computer processor and a computer memory operatively coupled to the computer processor, the computer memory having disposed within it computer program instructions capable of:

creating by a multimodal application a context, the context comprising words taken from at least one second web page different from the first web page, and/or taken from at least one grammar for the at least one second web page, wherein creating the context comprises maintaining a history of at least some words from previously visited web pages, generating occurrence counts for words in the history, and selecting words for the context based at least in part on the occurrence counts; and

supplementing by the multimodal application a grammar and/or lexicon for automatic speech recognition at the first web page in dependence upon the context.

13. The system of claim 12 , wherein creating the context further comprises inserting in the context one or more words from a grammar from a previously visited multimodal web page.

14. The system of claim 12 , wherein creating the context further comprises selecting words for the context from the history.

15. The system of claim 12 , wherein creating the context further comprises generating a frequent values list for words in the history, and selecting words for the context from the history in dependence upon the frequent values list.

16. The system of claim 12 , wherein creating the context further comprises generating a list of most recently used words in the history, and selecting words for the context from the history in dependence upon the list of most recently used words.

17. The system of claim 12 , wherein creating the context further comprises generating a list of words from the history used during a predetermined period of time, and selecting words for the context from the history in dependence upon the list of words from the history used during a predetermined period of time.

18. The system of claim 12 , wherein the grammar is a VoiceXML grammar.

19. The system of claim 12 , wherein the multimodal application runs on a voice server.

20. The system of claim 12 , wherein the computer program instructions are further capable of:

inferring at least one phoneme for at least one word in the grammar using text to speech conversion; and

adding the at least one phoneme for the at least one word to the lexicon.

21. The system of claim 12 , wherein the context comprises words taken from the at least one second web page.

22. The system of claim 12 , wherein the context comprises words taken from at least one grammar for the at least one second web page.

23. A computer program product for providing a context-based grammar for automatic speech recognition at a first web page, the computer program product comprising at least one recordable medium storing computer program instructions that, when executed, perform acts of:

creating by a multimodal application a context, the context comprising words taken from at least one second web page different from the first web page, and/or taken from at least one grammar for the at least one second web page, wherein creating the context comprises maintaining a history of at least some words from previously visited web pages, generating occurrence counts for words in the history, and selecting words for the context based at least in part on the occurrence counts; and

supplementing by the multimodal application a grammar and/or lexicon for automatic speech recognition at the first web page in dependence upon the context.

24. The computer program product of claim 23 , wherein creating the context further comprises inserting in the context one or more words from a grammar from a previously visited multimodal web page.

25. The computer program product of claim 23 , wherein creating the context further comprises selecting words for the context from the history.

26. The computer program product of claim 23 , wherein creating the context further comprises generating a frequent values list for words in the history, and selecting words for the context from the history in dependence upon the frequent values list.

27. The computer program product of claim 23 , wherein creating the context further comprises generating a list of most recently used words in the history, and selecting words for the context from the history in dependence upon the list of most recently used words.

28. The computer program product of claim 23 , wherein creating the context further comprises generating a list of words from the history used during a predetermined period of time, and selecting words for the context from the history in dependence upon the list of words from the history used during a predetermined period of time.

29. The computer program product of claim 23 , wherein the grammar is a VoiceXML grammar.

30. The computer program product of claim 23 , wherein the multimodal application runs on a voice server.

31. The computer program product of claim 23 , wherein the computer program instructions, when executed, further perform acts of:

inferring at least one phoneme for at least one word in the grammar using text to speech conversion; and

adding the at least one phoneme for the at least one word to the lexicon.

32. The computer program product of claim 23 , wherein the context comprises words taken from the at least one second web page.

33. The computer program product of claim 23 , wherein the context comprises words taken from at least one grammar for the at least one second web page.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065552/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →