IP Library Granted Patent US 8,135,591
Granted Patent B2
US 8,135,591 · App. 12/540,441 · Granted Mar 13, 2012

Method and system for training a text-to-speech synthesis system using a specific domain speech database

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,135,591
App. No.
12/540,441
Granted
Mar 13, 2012
Kind
B2
Abstract

A method and system are disclosed that train a text-to-speech synthesis system for use in speech synthesis. The method includes generating a speech database of audio files comprising domain-specific voices having various prosodies, and training a text-to-speech synthesis system using the speech database by selecting audio segments having a prosody based on at least one dialog state. The system includes a processor, a speech database of audio files, and modules for implementing the method.

Claims (35)

1. A system for training a text-to-speech synthesis system, the system comprising:

a processor;

a speech database of audio files comprising domain-specific voices having various prosodies; and

a module configured to control the processor to train a text-to-speech synthesis system using the speech database by selecting audio segments having a prosody based on at least one dialog state.

2. The system of claim 1 , wherein the dialog state is one of a beginning, middle and end of a dialog.

3. The system of claim 1 , further comprising a module configured to control the processor to train the text-to-speech synthesis system using the speech database by selecting audio segments having a prosody based on at least one speech act.

4. The system of claim 3 , wherein the speech act is based on at least one of syntax and semantics of a user's input.

5. The system of claim 1 , wherein the module configured to control the processor to train a text-to-speech synthesis system selects audio segments based on at least one of discourse and semantics.

6. The system of claim 1 , further comprising:

a module configured to control the processor to update the speech database using dialogs conducted with user's using the text-to-speech system.

7. The system of claim 1 , further comprising a module configured to control the processor to select audio segments based on a user's profile.

8. The system of claim 1 , further comprising:

a module configured to control the processor to tag audio files based on prosody for storage in the speech database.

9. A non-transitory computer-readable storage medium storing instructions for controlling a computing device for use in training a text-to-speech speech synthesis, the instructions comprising:

generating a speech database of audio files comprising domain-specific voices having various prosodies; and

training a text-to-speech synthesis system using the speech database by selecting audio segments having a prosody based on at least one dialog state.

10. The non-transitory computer-readable storage medium of claim 9 , wherein the dialog state is one of a beginning, middle and end of a dialog.

11. The non-transitory computer-readable storage medium of claim 9 , further comprising training the text-to-speech synthesis system using the speech database by selecting audio segments having a prosody based on at least one speech act.

12. The non-transitory computer-readable storage medium of claim 11 , wherein the speech act is based on at least one of syntax and semantics of a user's input.

13. The non-transitory computer-readable storage medium of claim 9 , further comprising selecting audio segments based on at least one of discourse and semantics.

14. The non-transitory computer-readable storage medium of claim 9 , further comprising:

updating the speech database using dialogs conducted with user's using the text-to-speech system.

15. The non-transitory computer-readable storage medium of claim 9 , further comprising selecting audio segments based on a user's profile.

16. The non-transitory computer-readable storage medium of claim 9 , further comprising:

tagging audio files based on prosody for storage in the speech database.

17. A method for training a text-to-speech speech synthesis system, the instructions comprising:

generating, via a processor, a speech database of audio files comprising domain-specific voices having various prosodies; and

training, via a processor, a text-to-speech synthesis system using the speech database by selecting audio segments having a prosody based on at least one dialog state.

18. The method of claim 17 , wherein the dialog state is one of a beginning, middle and end of a dialog.

19. The method of claim 17 , further comprising training the text-to-speech synthesis system using the speech database by selecting audio segments having a prosody based on at least one speech act.

20. The method of claim 19 , wherein the speech act is based on at least one of syntax and semantics of a user's input.

21. The method of claim 17 , further comprising:

updating the speech database using dialogs conducted with user's using the text-to-speech system.

22. The method of claim 17 , further comprising:

tagging audio files based on prosody for storage in the speech database.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2017
From: SCHROETER, HORST JUERGEN
To: AT&T CORP.
Reel/Frame 041047/0245 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2017
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 041047/0519 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNOR/ASSIGNEE NAME PREVIOUSLY RECORDED ON REEL 025152 FRAME 0243. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNOR NAME SHOULD READ ONLY: AT&T CORP. ASSIGNEE NAME SHOULD READ ONLY: AT&T PROPERTIES, LLC. Recorded Jan 23, 2017
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 041072/0258 →