IP Library Granted Patent US 9,626,695
Granted Patent B2
US 9,626,695 · App. 14/451,151 · Granted Apr 18, 2017

Automatically presenting different user experiences, such as customized voices in automated communication systems

Inventors: Sundar Balasubramanian (Seattle, WA); Michael McSherry (Seattle, WA); Eric Jun Fu (Bellevue, WA); Daniel Hendrick (Bellevue, WA); Deepankar Katyal (Seattle, WA); David J. Kay (Seattle, WA)
Assignee: NUANCE COMMUNICATIONS, INC.
G06Q30/0255G06Q30/0257G06Q30/0261G06Q30/0267G06Q30/0269G10L13/033G10L15/20H04L67/2847
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,626,695
App. No.
14/451,151
Granted
Apr 18, 2017
Kind
B2
Abstract

An automated communication system with an associated method for presenting customized voices is disclosed. The system which performs a predetermined task accepts information regarding an intended user indicating the intended user's identity, preferences, etc. Next, the system customizes one or more voices for the intended user based on the accepted information. The system then presents to the intended user one or more audible communications converted from text associated with a predetermined task performed by the system using the one or more customized voices.

Claims (72)

1. At least one non-transitory computer-readable medium, carrying instructions, which when executed by at least one hardware data processor, performs a method of presenting customized voices by an automated communication system, the method comprising:

accepting information regarding an intended user of the automated communication system, wherein accepting the information includes:

i) analyzing text and voice input from the user, and

ii) deducing an interaction style of the user based on the text and voice input from the user,

wherein the user's interaction style indicates a mood of the user and an interaction style that the user desires, and

wherein the mood includes being available or busy and wherein the interaction style includes a formal or an informal style, and

iii) storing data regarding the user's interaction style;

customizing or selecting one of multiple text-to-speech voices based on the accepted information,

wherein the multiple voices are audio properties of the multiple voices that vary by volume and pitch, and

wherein the multiple voices include an irreverent voice and a succinct voice; and

presenting, to the user and for audible output to the user on a computing device of the user, one or more audible communications for a task performed by the automated communication system using the one or more customized or selected voices.

2. The non-transitory computer-readable medium of claim 1 , wherein the automated communication system is a server, and wherein the task is a computerized event invitation provided by the server to the computing device of the user, a computerized greeting card provided by the server to the computing device of the user, a computerized game provided by the server to the computing device of the user, or a computerized storybook provided by the server to the computing device of the user, and wherein the server computer also provides to the user, via the computing device, a spoken advertisement.

3. The non-transitory computer-readable medium of claim 1 , further comprising

storing multiple predetermined voices,

wherein customizing or selecting the one of multiple voices includes selecting the one voice from the stored voices, and wherein the voices are each different computer-generated voices.

4. The non-transitory computer-readable medium of claim 1 , further comprising

storing multiple predetermined voice components,

wherein the customizing or selecting the one of multiple voices includes synthesizing the one voice using one or more of the stored voice components.

5. The non-transitory computer-readable medium of claim 1 , wherein

the intended user interacts with the automated communication system, and

the one or more audible communications are presented in response to inquiries from the user, and

wherein the method further comprises customizing or selecting the one or more voices based on a nature of one of the inquiries or a state of the task.

6. The non-transitory computer-readable medium of claim 1 , wherein the one voice is selected based on an occasion and an intended audience to receive the one or more audible communications, and wherein the task is an electronic invitation to attend a child's party, to attend a woman's wedding, or to attend a professor's retirement party.

7. The non-transitory computer-readable medium of claim 1 , wherein the one or more audible communications correspond to text associated with the task, wherein the task is a game that uses multiple, different voices to present different questions depending on contestants playing the game, depending on question categories, and depending on prizes involved.

8. The non-transitory computer-readable medium of claim 1 , wherein the accepted information indicates a profession or a hobby of the user.

9. The non-transitory computer-readable medium of claim 1 , further comprising:

identifying multiple stages of or roles in the task; and

assigning different customized or selected voices to the stages or roles,

wherein the presenting is performed based on a current stage of or an active role in the task.

10. A method of operating an automated communication system to present customized voices, the method comprising:

accepting information regarding an intended user of the automated communication system, wherein accepting the information includes:

i) analyzing text and voice input from the user, and

ii) deducing an interaction style of the user based on the text and voice input from the user,

wherein the user's interaction style indicates a mood of the user and an interaction style that the user desires, and

wherein the mood includes being available or busy and wherein the interaction style includes a formal or an informal style, and

iii) storing data regarding the user's interaction style;

customizing or selecting one of multiple text-to-speech voices based on the accepted information,

wherein the multiple voices are audio properties of the multiple voices that vary by volume and pitch,

wherein the customizing or selecting is performed by a hardware processor, and

wherein the multiple voices include an irreverent voice and a succinct voice; and

presenting, to the user and for audible output to the user on a computing device of the user, one or more audible communications for a task performed by the automated communication system using the one or more customized or selected voices.

11. The method of claim 10 , wherein the automated communication system is a server, and wherein the task is a computerized event invitation provided by the server to the computing device of the user, a computerized greeting card provided by the server to the computing device of the user, a computerized game provided by the server to the computing device of the user, or a computerized storybook provided by the server to the computing device of the user, and wherein the server computer also provides to the user, via the computing device, a spoken advertisement.

12. The method of claim 10 , further comprising

storing multiple predetermined voices,

wherein customizing or selecting the one of multiple voices includes selecting the one voice from the stored voices, and wherein the voices are each different computer-generated voices.

13. The method of claim 10 , further comprising

storing multiple predetermined voice components,

wherein the customizing or selecting the one of multiple voices includes synthesizing the one voice using one or more of the stored voice components.

14. The method of claim 10 , wherein

the intended user interacts with the automated communication system, and

the one or more audible communications are presented in response to inquiries from the user, and

wherein the method further comprises customizing or selecting the one or more voices based on a nature of one of the inquiries or a state of the task.

15. The method of claim 10 , wherein the one voice is selected based on an occasion and an intended audience to receive the one or more audible communications, and wherein the task is an electronic invitation to attend a child's party, to attend a woman's wedding, or to attend a professor's retirement party.

16. The method of claim 10 , wherein the one or more audible communications correspond to text associated with the task, wherein the task is a game that uses multiple, different voices to present different questions depending on contestants playing the game, depending on question categories, and depending on prizes involved.

17. The method of claim 10 , wherein the accepted information indicates a profession or a hobby of the user.

18. The method of claim 10 , further comprising:

identifying multiple stages of or roles in the task; and

assigning different customized or selected voices to the stages or roles,

wherein the presenting is performed based on a current stage of or an active role in the task.

19. An automated communication system for presenting customized voices, the system comprising:

a processor configured to:

accept information regarding an intended user of the automated communication system, wherein accepting the information includes:

i) analyzing text and voice input from the user, and

ii) deducing an interaction style of the user based on the text and voice input from the user,

wherein the user's interaction style indicates a mood of the user and an interaction style that the user desires, and

wherein the mood includes being available or busy and wherein the interaction style includes a formal or an informal style;

customize or select one of multiple text-to-speech voices based on the accepted information,

wherein the multiple voices are audio properties of the multiple voices that vary by volume and pitch, and

wherein the multiple voices include an irreverent voice and a succinct voice; and

present, for the user and for audible output to the user on a computing device of the user, one or more audible communications for a task performed by the automated communication system using the one or more customized or selected voices; and

memory, coupled to the processor, configured to store data regarding the user's interaction style.

20. The system of claim 19 , wherein the processor and the memory are part of a server, and wherein the task is a computerized event invitation provided by the server to the computing device of the user, a computerized greeting card provided by the server to the computing device of the user, a computerized game provided by the server to the computing device of the user, or a computerized storybook provided by the server to the computing device of the user, and wherein the server computer also provides to the user, via the computing device, a spoken advertisement.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065552/0934 →
Continuity (2)
Continuation 14316663 · Jun 26, 2014
Related Publication 20150379981A1 · Dec 31, 2015