IP Library Granted Patent US 7,584,104
Granted Patent B2
US 7,584,104 · App. 11/530,258 · Granted Sep 1, 2009

Method and system for training a text-to-speech synthesis system using a domain-specific speech database

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,584,104
App. No.
11/530,258
Granted
Sep 1, 2009
Kind
B2
Abstract

A system, method and computer readable medium that trains a text-to-speech synthesis system for use in speech synthesis is disclosed. The method may include recording audio files of one or more live voices speaking language used in a specific domain, the audio files being recorded using various prosodies, storing the recorded audio files in a speech database; and training a text-to-speech synthesis system using the speech database, wherein the text-to-speech synthesis system selects audio selects audio segments having a prosody based on at least one dialog state and one speech act.

Claims (32)

1. A method for training a text-to-speech synthesis system, comprising:

recording audio files of one or more live voices speaking language used in a specific domain, the audio files being recorded using various prosodies;

storing the recorded audio files in a speech database; and

training a text-to-speech synthesis system using the speech database, wherein the text-to-speech synthesis system selects audio segments having a prosody based on at least one dialog state and one speech act.

2. The method of claim 1 , wherein the dialog state is one of a beginning, middle and end of a dialog.

3. The method of claim 1 , wherein the speech act is based on at least one of syntax and semantics of a user's input.

4. The method of claim 1 , wherein the text-to-speech synthesis system selects audio segments based on at least one of discourse and semantics.

5. The method of claim 1 , further comprising:

updating the speech database using dialogs conducted with user's using the text-to-speech system.

6. The method of claim 1 , wherein the text-to-speech synthesis system selects audio segments based on a user's profile.

7. The method of claim 1 , further comprising:

tagging audio files based on prosody for storage in the speech database.

8. A computer-readable storage medium storing instructions for controlling a computing device for use in training a text-to-speech speech synthesis, the instructions comprising:

recording audio files of one or more live voices speaking language used in a specific domain, the audio files being recorded using various prosodies;

storing the recorded audio files in a speech database; and

training a text-to-speech synthesis system using the speech database, wherein the text-to-speech synthesis system selects audio segments having a prosody based on at least one dialog state and one speech act.

9. The computer-readable storage medium of claim 8 , wherein the dialog state is one of a beginning, middle and end of a dialog.

10. The computer-readable storage medium of claim 8 , wherein the speech act is based on at least one of syntax and semantics of a user's input.

11. The computer-readable storage medium of claim 8 , wherein the text-to-speech synthesis system selects audio segments based on at least one of discourse and semantics.

12. The computer-readable storage medium of claim 8 , further comprising:

updating the speech database using dialogs conducted with user's using the text-to-speech system.

13. The computer-readable storage medium of claim 8 , wherein the text-to-speech synthesis system selects audio segments based on a user's profile.

14. The computer-readable storage medium of claim 8 , further comprising:

tagging audio files based on prosody for storage in the speech database.

15. A text-to-speech synthesis training system, comprising:

a speech database; and

a specific domain speech knowledge module that record audio files of one or more live voices speaking language used in a specific domain, the audio files being recorded using various prosodies, stores the recorded audio files in the speech database, and trains a text-to-speech synthesis system using the speech database, wherein the text-to-speech synthesis system selects audio segments having a prosody based on at least one dialog state and one speech act.

16. The system of claim 15 , wherein the dialog state is one of a beginning, middle and end of a dialog.

17. The system of claim 15 , wherein the speech act is based on at least one of syntax and semantics of a user's input.

18. The system of claim 15 , wherein the text-to-speech synthesis system selects audio files based on at least one of discourse and semantics.

19. The system of claim 15 , wherein the specific domain speech knowledge module updates the speech database using dialogs conducted with user's using the text-to-speech system.

20. The system of claim 15 , wherein the text-to-speech synthesis system selects audio segments based on a user's profile.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065533/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038275/0238 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038275/0310 →