IP Library Patent Application 11690471
Patent Application
App. No. 11/690,471

Speech-Enabled Predictive Text Selection For A Multimodal Application

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
11/690,471
Abstract

Methods, apparatus, and products are disclosed for speech-enabled predictive text selection for a multimodal application, the multimodal application operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to an automatic speech recognition (‘ASR’) engine through a VoiceXML interpreter, including: identifying, by the VoiceXML interpreter, a text prediction event, the text prediction event characterized by one or more predictive texts for a text input field of the multimodal application; creating, by the VoiceXML interpreter, a grammar in dependence upon the predictive texts; receiving, by the VoiceXML interpreter, a voice utterance from a user; and determining, by the VoiceXML interpreter using the ASR engine, recognition results in dependence upon the voice utterance and the grammar, the recognition results representing a user selection of a particular predictive text.

Claims (41)

1 . A computer-implemented method of speech-enabled predictive text selection for a multimodal application, the multimodal application operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to an automatic speech recognition (‘ASR’) engine through a VoiceXML interpreter, the method comprising:

identifying, by the VoiceXML interpreter, a text prediction event, the text prediction event characterized by one or more predictive texts for a text input field of the multimodal application;

creating, by the VoiceXML interpreter, a grammar in dependence upon the predictive texts;

receiving, by the VoiceXML interpreter, a voice utterance from a user; and

determining, by the VoiceXML interpreter using the ASR engine, recognition results in dependence upon the voice utterance and the grammar, the recognition results representing a user selection of a particular predictive text.

2 . The method of claim 1 further comprising rendering, by the VoiceXML interpreter, at least a portion of the recognition results in the text input field.

3 . The method of claim 1 further comprising:

creating, by the VoiceXML interpreter, a user prompt for the voice utterance in dependence upon the predictive texts; and

prompting, by the VoiceXML interpreter, the user for the voice utterance in dependence upon the user prompt.

4 . The method of claim 1 further comprising rendering, by a multimodal browser, the predictive texts on a graphical user interface of the multimodal device in dependence upon the text prediction event.

5 . The method of claim 1 wherein creating, by the VoiceXML interpreter, a grammar in dependence upon the predictive texts further comprises:

generating a grammar rule for the grammar, the grammar rule specifying each predictive text as an alternative for recognition.

6 . The method of claim 1 wherein the text prediction event occurs when the user types a character in the text input field of the multimodal application.

7 . The method of claim 1 wherein the text prediction event occurs when the user speaks a character for input in the text input field of the multimodal application.

8 . Apparatus for speech-enabled predictive text selection for a multimodal application, the multimodal application operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to an automatic speech recognition (‘ASR’) engine through a VoiceXML interpreter, the apparatus comprising a computer processor and a computer memory operatively coupled to the computer processor, the computer memory having disposed within it computer program instructions capable of:

identifying, by the VoiceXML interpreter, a text prediction event, the text prediction event characterized by one or more predictive texts for a text input field of the multimodal application;

creating, by the VoiceXML interpreter, a grammar in dependence upon the predictive texts;

receiving, by the VoiceXML interpreter, a voice utterance from a user; and

determining, by the VoiceXML interpreter using the ASR engine, recognition results in dependence upon the voice utterance and the grammar, the recognition results representing a user selection of a particular predictive text.

9 . The apparatus of claim 8 further comprising computer program instructions capable of rendering, by the VoiceXML interpreter, at least a portion of the recognition results in the text input field.

10 . The apparatus of claim 8 further comprising computer program instructions capable of:

creating, by the VoiceXML interpreter, a user prompt for the voice utterance in dependence upon the predictive texts; and

prompting, by the VoiceXML interpreter, the user for the voice utterance in dependence upon the user prompt.

11 . The apparatus of claim 8 further comprising computer program instructions capable of rendering, by a multimodal browser, the predictive texts on a graphical user interface of the multimodal device in dependence upon the text prediction event.

12 . The apparatus of claim 8 wherein creating, by the VoiceXML interpreter, a grammar in dependence upon the predictive texts further comprises:

generating a grammar rule for the grammar, the grammar rule specifying each predictive text as an alternative for recognition.

13 . The apparatus of claim 8 wherein the text prediction event occurs when the user speaks a character for input in the text input field of the multimodal application.

14 . A computer program product for speech-enabled predictive text selection for a multimodal application, the multimodal application operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to an automatic speech recognition (‘ASR’) engine through a VoiceXML interpreter, the computer program product disposed upon a computer-readable medium, the computer program product comprising computer program instructions capable of:

identifying, by the VoiceXML interpreter, a text prediction event, the text prediction event characterized by one or more predictive texts for a text input field of the multimodal application;

creating, by the VoiceXML interpreter, a grammar in dependence upon the predictive texts;

receiving, by the VoiceXML interpreter, a voice utterance from a user; and

determining, by the VoiceXML interpreter using the ASR engine, recognition results in dependence upon the voice utterance and the grammar, the recognition results representing a user selection of a particular predictive text.

15 . The computer program product of claim 14 further comprising computer program instructions capable of rendering, by the VoiceXML interpreter, at least a portion of the recognition results in the text input field.

16 . The computer program product of claim 14 further comprising computer program instructions capable of:

creating, by the VoiceXML interpreter, a user prompt for the voice utterance in dependence upon the predictive texts; and

prompting, by the VoiceXML interpreter, the user for the voice utterance in dependence upon the user prompt.

17 . The computer program product of claim 14 further comprising computer program instructions capable of rendering, by a multimodal browser, the predictive texts on a graphical user interface of the multimodal device in dependence upon the text prediction event.

18 . The computer program product of claim 14 wherein creating, by the VoiceXML interpreter, a grammar in dependence upon the predictive texts further comprises:

generating a grammar rule for the grammar, the grammar rule specifying each predictive text as an alternative for recognition.

19 . The computer program product of claim 14 wherein the text prediction event occurs when the user types a character in the text input field of the multimodal application.

20 . The computer program product of claim 14 wherein the text prediction event occurs when the user speaks a character for input in the text input field of the multimodal application.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2007
From: CROSS, CHARLES W.; JABLOKOV, IGOR R.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 019879/0185 →