IP Library Granted Patent US 8,725,513
Granted Patent B2
US 8,725,513 · App. 11/734,422 · Granted May 13, 2014

Providing expressive user interaction with a multimodal application

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,725,513
App. No.
11/734,422
Granted
May 13, 2014
Kind
B2
Abstract

Methods, apparatus, and products are disclosed for providing expressive user interaction with a multimodal application, the multimodal application operating in a multimodal browser on a multimodal device supporting multiple modes of user interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to a speech engine through a VoiceXML interpreter, including: receiving, by the multimodal browser, user input from a user through a particular mode of user interaction; determining, by the multimodal browser, user output for the user in dependence upon the user input; determining, by the multimodal browser, a style for the user output in dependence upon the user input, the style specifying expressive output characteristics for at least one other mode of user interaction; and rendering, by the multimodal browser, the user output in dependence upon the style.

Claims (84)

1. A computer-implemented method of providing expressive user interaction with a multimodal application, the multimodal application operating in a multimodal browser on a multimodal device supporting multiple modes of user interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to a speech engine through a VoiceXML interpreter, the method comprising:

receiving, by the multimodal browser, user input from a user through a particular mode of user interaction;

determining, by the multimodal browser, user output for the user in dependence upon the user input;

determining, by the multimodal browser, a style for the user output in dependence upon the user input, the style specifying expressive output characteristics for at least one other mode of user interaction; and

rendering, by the multimodal browser, the user output in dependence upon the style,

wherein determining the style for the user output in dependence upon the user input comprises performing a determination distinct from determining the user output,

wherein determining, by the multimodal browser, a style for the user output in dependence upon the user input comprises determining, by the multimodal browser, a style for the user output in dependence upon meaning of the user input.

2. The method of claim 1 further comprising prompting, by the multimodal browser, the user for the user input.

3. The method of claim 1 wherein:

receiving, by the multimodal browser, user input from the user through the particular mode of user interaction further comprises receiving graphical user input from the user through a graphical user interface;

determining, by the multimodal browser, user output for the user in dependence upon the user input further comprises determining the user output for the user in dependence upon the graphical user input; and

determining, by the multimodal browser, the style for the user output further comprises determining the style for the user output in dependence upon the graphical user input.

4. The method of claim 1 wherein:

receiving, by the multimodal browser, user input from the user through the particular mode of user interaction further comprises:

receiving a voice utterance from the user, and

determining recognition results using the speech engine and a grammar;

determining, by the multimodal browser, user output for the user in dependence upon the user input further comprises determining the user output for the user in dependence upon the recognition results; and

determining, by the multimodal browser, the style for the user output further comprises determining the style for the user output in dependence upon the recognition results.

5. The method of claim 1 wherein:

the style specifies prosody for the voice mode of user interaction; and

rendering, by the multimodal browser, the user output in dependence upon the style further comprises:

synthesizing, through the VoiceXML interpreter using the speech engine, the user output into synthesized speech in dependence upon the style, and

playing the synthesized speech for the user.

6. The method of claim 5 further comprising providing, by the speech engine to the multimodal browser through the VoiceXML interpreter, a prosody event in dependence upon the synthesizing of the user output into synthesized speech.

7. The method of claim 1 wherein:

the style specifies visual characteristics for a visual mode of user interaction; and

rendering, by the multimodal browser, the user output in dependence upon the style further comprises displaying the user output to the user in dependence upon the style.

8. The method of claim 7 further comprising receiving, by the multimodal browser through the VoiceXML interpreter, a prosody event from the speech engine, and wherein rendering, by the multimodal browser, the user output in dependence upon the style further comprises rendering the user output in dependence upon the prosody event.

9. Apparatus for providing expressive user interaction with a multimodal application, the multimodal application operating in a multimodal browser on a multimodal device supporting multiple modes of user interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to a speech engine through a VoiceXML interpreter, the apparatus comprising a computer processor and a computer memory operatively coupled to the computer processor, the computer memory having stored within it computer program instructions which, when executed by the computer processor, cause performance of a method comprising:

receiving, by the multimodal browser, user input from a user through a particular mode of user interaction;

determining, by the multimodal browser, user output for the user in dependence upon the user input;

determining, by the multimodal browser, a style for the user output in dependence upon the user input, the style specifying expressive output characteristics for at least one other mode of user interaction; and

rendering, by the multimodal browser, the user output in dependence upon the style,

wherein determining the style for the user output in dependence upon the user input comprises performing a determination distinct from determining the user output,

wherein determining, by the multimodal browser, a style for the user output in dependence upon the user input comprises determining, by the multimodal browser, a style for the user output in dependence upon meaning of the user input.

10. The apparatus of claim 9 , wherein the computer memory also has stored within it computer program instructions which, when executed by the computer processor, cause prompting, by the multimodal browser, the user for the user input.

11. The apparatus of claim 9 wherein:

receiving, by the multimodal browser, user input from the user through the particular mode of user interaction further comprises receiving graphical user input from the user through a graphical user interface;

determining, by the multimodal browser, user output for the user in dependence upon the user input further comprises determining the user output for the user in dependence upon the graphical user input; and

determining, by the multimodal browser, the style for the user output further comprises determining the style for the user output in dependence upon the graphical user input.

12. The apparatus of claim 9 wherein:

receiving, by the multimodal browser, user input from the user through the particular mode of user interaction further comprises:

receiving a voice utterance from the user, and

determining recognition results using the speech engine and a grammar;

determining, by the multimodal browser, user output for the user in dependence upon the user input further comprises determining the user output for the user in dependence upon the recognition results; and

determining, by the multimodal browser, the style for the user output further comprises determining the style for the user output in dependence upon the recognition results.

13. The apparatus of claim 9 wherein:

the style specifies prosody for the voice mode of user interaction;

rendering, by the multimodal browser, the user output in dependence upon the style further comprises:

synthesizing, through the VoiceXML interpreter using the speech engine, the user output into synthesized speech in dependence upon the style, and

playing the synthesized speech for the user; and

the computer memory also has stored within it computer program instructions which, when executed by the computer processor, cause providing, by the speech engine to the multimodal browser through the VoiceXML interpreter, a prosody event in dependence upon the synthesizing of the user output into synthesized speech.

14. The apparatus of claim 9 wherein:

the style specifies visual characteristics for a visual mode of user interaction;

the computer memory also has stored within it computer program instructions which, when executed by the computer processor, cause receiving, by the multimodal browser through the VoiceXML interpreter, a prosody event from the speech engine; and

rendering, by the multimodal browser, the user output in dependence upon the style further comprises displaying the user output to the user in dependence upon the style and rendering the user output in dependence upon the prosody event.

15. A computer-recordable device storing computer program instructions for providing expressive user interaction with a multimodal application, the multimodal application operating in a multimodal browser on a multimodal device supporting multiple modes of user interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to a speech engine through a VoiceXML interpreter, wherein the computer program instructions, when executed, cause performance of a method comprising:

receiving, by the multimodal browser, user input from a user through a particular mode of user interaction;

determining, by the multimodal browser, user output for the user in dependence upon the user input;

determining, by the multimodal browser, a style for the user output in dependence upon the user input, the style specifying expressive output characteristics for at least one other mode of user interaction; and

rendering, by the multimodal browser, the user output in dependence upon the style,

wherein determining the style for the user output in dependence upon the user input comprises performing a determination distinct from determining the user output,

wherein determining, by the multimodal browser, a style for the user output in dependence upon the user input comprises determining, by the multimodal browser, a style for the user output in dependence upon meaning of the user input.

16. The computer-recordable device of claim 15 , wherein the method further comprises prompting, by the multimodal browser, the user for the user input.

17. The computer-recordable device of claim 15 wherein:

receiving, by the multimodal browser, user input from the user through the particular mode of user interaction further comprises receiving graphical user input from the user through a graphical user interface;

determining, by the multimodal browser, user output for the user in dependence upon the user input further comprises determining the user output for the user in dependence upon the graphical user input; and

determining, by the multimodal browser, the style for the user output further comprises determining the style for the user output in dependence upon the graphical user input.

18. The computer-recordable device of claim 15 wherein:

receiving, by the multimodal browser, user input from the user through the particular mode of user interaction further comprises:

receiving a voice utterance from the user, and

determining recognition results using the speech engine and a grammar;

determining, by the multimodal browser, user output for the user in dependence upon the user input further comprises determining the user output for the user in dependence upon the recognition results; and

determining, by the multimodal browser, the style for the user output further comprises determining the style for the user output in dependence upon the recognition results.

19. The computer-recordable device of claim 15 wherein:

the style specifies prosody for the voice mode of user interaction;

rendering, by the multimodal browser, the user output in dependence upon the style further comprises:

synthesizing, through the VoiceXML interpreter using the speech engine, the user output into synthesized speech in dependence upon the style, and

playing the synthesized speech for the user; and

the method further comprises providing, by the speech engine to the multimodal browser through the VoiceXML interpreter, a prosody event in dependence upon the synthesizing of the user output into synthesized speech.

20. The computer-recordable device of claim 15 wherein:

the style specifies visual characteristics for a visual mode of user interaction;

the method further comprising receiving, by the multimodal browser through the VoiceXML interpreter, a prosody event from the speech engine; and

rendering, by the multimodal browser, the user output in dependence upon the style further comprises displaying the user output to the user in dependence upon the style and rendering the user output in dependence upon the prosody event.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →