IP Library Granted Patent US 8,407,056
Granted Patent B2
US 8,407,056 · App. 13/179,294 · Granted Mar 26, 2013

Global speech user interface

Inventors: Adam Jordan (Oakland, CA); Scott Lynn Maddux (San Francisco, CA); Tim Plowman (Berkeley, CA); Victoria Stanbach (Redwood City, CA); Jody Williams (San Carlos, CA)
Assignee: Promptu Systems Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,407,056
App. No.
13/179,294
Granted
Mar 26, 2013
Kind
B2
Abstract

A global speech user interface (GSUI) comprises an input system to receive a user's spoken command, a feedback system along with a set of feedback overlays to give the user information on the progress of his spoken requests, a set of visual cues on the television screen to help the user understand what he can say, a help system, and a model for navigation among applications. The interface is extensible to make it easy to add new applications.

Claims (83)

1. A non-transitory computer readable storage medium encoded with instructions, which when executed by a communications system, configures a processor to establish a global speech user interface (GSUI) system comprising modules for:

performing speech recognition to transcribe spoken commands into commands acceptable by said communications system;

using the transcribed spoken commands to navigate among applications hosted on said communications system; and

displaying a set of visual cues to guide a user in issuing proper spoken commands, said visual cues comprising:

a set of immediate speech feedback overlays, each of which provides non-textual feedback information about a state of said communications system;

a set of help overlays, each of which provides a context-sensitive list of frequently used speech-activated commands for each screen of every speech-activated application;

a set of feedback overlays, each of which provides information about a problem that said communications system is experiencing; and

a main menu overlay that shows a list of services available to the user, each of said services being accessible by spoken command;

further comprising a user center that provides any of:

training and tutorials on how to use said communications system;

more help with specific speech-activated applications;

user account management; and

user settings and preferences for said communications system;

wherein said set of feedback overlays comprises any of:

a set of recognition feedback overlays that informs the user of a situation related to recognition; and

a set of application overlays that informs the user of an error or a problem related to an application used in said GSUI; and

wherein said set of recognition feedback overlays, in responding to unsuccessful recognitions that immediately follow one another, is displayed in three different modes comprising:

a first mode wherein said immediate speech feedback indicator changes to a question mark in responding to the first unsuccessful recognition;

a second mode wherein a textual message and a link to said help overlay are displayed in responding to the second unsuccessful recognition; and

a third mode wherein a textual message, a link to said help overlay, and a link to said more help overlay are displayed in responding to the third and subsequent unsuccessful recognition.

2. The medium of claim 1 , wherein each of said immediate speech feedback overlays provides non-textual feedback information about a state of said communications system, said state being any of:

listening to the user's spoken command;

non-speech enabled alert;

speech recognition processing;

application alert;

positive speech recognition; and

speech recognition unsuccessful.

3. The medium of claim 1 , wherein each of said help overlays is accessible at all times.

4. The medium of claim 1 , wherein said list of speech-activated commands provided by said help overlay comprises any of:

a set of application-specific commands;

a command associated with the user center for more help;

a command associated to said main menu display; and

a command to make said overlay disappear.

5. The medium of claim 1 , wherein said visual cues further comprise a treatment of on-screen text which can be activated by a spoken command.

6. The medium of claim 5 , wherein said treatment is an overlay in round shape and green color.

7. The medium of claim 5 , wherein said treatment can be turned on or off by the user.

8. The medium of claim 5 , wherein said on-screen text comprises any of:

a static text used in labels for on-screen graphics or in virtual buttons that may be selected by a cursor; and

a dynamic text used in content wherein one or more words can be activated by a spoken command.

9. The medium of claim 1 , wherein any of said help overlays, feedback overlays and main menu overlay is implemented in a dialog box, said dialog box comprising any of:

one or more text box for textual information; and

one or more virtual buttons.

10. The medium of claim 9 , wherein said dialog box further comprises an identity indicator.

11. The medium of claim 10 , wherein said dialog box has either of an approximately transparent background, and an opaque background.

12. The medium of claim 11 , wherein said approximately transparent background is incorporated with any of a dynamic image to enhance said identity indicator and a static image to enhance said identity indicator.

13. The medium of claim 11 , wherein said text box is overlaid on said approximately transparent background.

14. The medium of claim 1 , further comprising a speaker personalization and identification mechanism that allows users to train said communications system to identify respective users by voice.

15. The medium of claim 1 , further including a remote control to relay user speech directly or indirectly for performing speech recognition, the remote control including a push-to-talk button;

where the help overlays include a volume indicator in the same color as the push-to-talk button and on-screen graphics indicating speech-enabled user interface elements.

16. The medium of claim 1 , wherein said displaying provides immediate real-time visual feedback indicating various states of speech recognition activities.

17. The medium of claim 16 , said real-time visual feedback comprises a set of overlays, each of which provides non-textual feedback information indicating state information selected from a predetermined group speech recognition activity states, the group including all of:

receiving spoken utterance;

processing utterance;

successful recognition;

unsuccessful recognition; and

command not allowed.

18. The medium of claim 1 , wherein said displaying allows the user to initiate, via spoken command, an overlay display which indicates selectable user interface elements.

19. The medium of claim 18 , wherein said selectable user interface elements comprise any of:

numeric identifications;

navigation options; and

application control options.

20. The medium of claim 1 , wherein when the user's spoken command is not recognized with a predefined degree of confidence, said displaying presents a list of predicted commands prompting the user to select from said list.

21. The medium of claim 1 , further comprising:

navigating one or more on-screen lists based information via spoken commands.

22. The medium of claim 21 , wherein said navigating enables the user to perform any of:

direct said on-screen list based information scroll up or scroll down by speaking a corresponding command;

select an item from said on-screen list based information by speaking a letter or a number identifying said item; and

select an item from said on-screen list based information by speaking the name of said item.

23. The medium of claim 1 , further comprising any of:

allowing the user to navigate directly between applications via spoken command or a speech enabled menu; and

allowing the user to navigate directly to previously book-marked pages via spoken command.

24. The medium of claim 23 , wherein said direct navigation to previously book-marked pages operates within and between applications.

25. The medium of claim 1 , further including providing an interactive program guide that the user can access via spoken command.

26. The medium of claim 1 , further including providing an interactive video on demand service, from which the user can order any video program contained in a list.

27. The medium of claim 26 , wherein said video on demand service comprises any of:

allowing the user to, via spoken command, sort video programs by categories;

allowing the user to, via spoken command, search video programs by properties;

allowing the user to, via spoken command, set parental control with which children are blocked from accessing controlled video programs; and

allowing the user to obtain automatic recommendation based on voice identification.

28. The medium of claim 1 , further including:

allowing the user to complete all aspects of a transaction via spoken commands.

29. The medium of claim 1 , further including:

allowing the user to exercise control, via spoken commands, over home services and devices.

Continuity (4)
Continuation 11933191 · Oct 31, 2007
Division 10260906 · Sep 30, 2002
Provisional Application 60327207 · Oct 3, 2001
Related Publication 20110270615A1 · Nov 3, 2011