IP Library Granted Patent US 7,630,898
Granted Patent B1
US 7,630,898 · App. 11/235,817 · Granted Dec 8, 2009

System and method for preparing a pronunciation dictionary for a text-to-speech voice

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,630,898
App. No.
11/235,817
Granted
Dec 8, 2009
Kind
B1
Abstract

Disclosed are various elements of a toolkit used for generating a TTS voice for use in a spoken dialog system. The embodiments in each case may be in the form of the system, a computer-readable medium or a method for generating the TTS voice. One embodiment of the invention relates to a method of generating a database for a TTS voice. The method comprises matching every spoken word associated with a TTS voice database with a smallest set of possible pronunciations for each word. The smallest set is generated by automatically determining a dialect and linguistic context using linguistic rules, empirically determining idiosyncratic speaker characteristics and determining a subject domain. The method further comprises dynamically generating a pronunciation dictionary on a word-by-word basis using the smallest set.

Claims (44)

1. A computer-implemented method of generating a database for a text-to-speech (TTS) voice, the method comprising:

matching via a processor every spoken word associated with a TTS voice database with a smallest set of possible pronunciations for each word, the smallest set being generated by:

automatically via the processor determining a dialect and linguistic context using linguistic rules;

empirically determining idiosyncratic speaker characteristics; and

determining a subject domain; and

dynamically generating a pronunciation dictionary on a word-by-word basis using the smallest set.

2. The computer-implemented method of claim 1 , wherein the dynamically generated pronunciation dictionary accounts for reading errors and idiosyncrasies of speakers.

3. The computer-implemented method of claim 1 , further comprising:

forcing an automatic speech recognition (ASR) module to choose from a subset of one or more variants of a word when more than one pronunciation variants exists for a given word.

4. The computer-implemented method of claim 1 , further comprising:

automatically generating phonetic variant pronunciations for the pronunciation dictionary for any given word.

5. The computer-implemented method of claim 4 , wherein generating phonetic variant pronunciations is based on a surrounding linguistic context for each word.

6. The computer-implemented method of claim 1 , wherein the linguistic contexts are associated with foreign languages.

7. The computer-implemented method of claim 1 , further comprising:

tracking whether each lexical pronunciation is either machine generated or human entered.

8. The computer-implemented method of claim 7 , further comprising:

flagging machine generated lexical pronunciations for human inspection.

9. The computer-implemented method of claim 1 , further comprising adding default pronunciations to the pronunciation dictionary based on TTS letter-to-sound rules.

10. A computing device for generating a database for a text-to-speech (TTS) voice, the computing device comprising:

a module configured to control the processor to match every spoken word associated with a TTS voice database with a smallest set of possible pronunciations for each word, the smallest set being generated by:

automatically via the processor determining a dialect and linguistic context using linguistic rules;

empirically determining idiosyncratic speaker characteristics; and

determining a subject domain; and

a module configured to control the processor to dynamically generate a pronunciation dictionary on a word-by-word basis using the smallest set.

11. The computing device of claim 10 , wherein the dynamically generated pronunciation dictionary accounts for reading errors and idiosyncrasies of speakers.

12. The computing device of claim 10 , further comprising:

a module configured to force an automatic speech recognition (ASR) module to choose from a subset of one or more variants of a word when more than one pronunciation variants exist for a given word.

13. The computing device of claim 10 , further comprising:

a module configured to automatically generate phonetic variant pronunciations for the pronunciation dictionary for any given word.

14. The computing device of claim 13 , wherein generating phonetic variant pronunciations is based on a surrounding linguistic context for each word.

15. The computing device of claim 9 , wherein the linguistic contexts are associated with foreign languages.

16. The computing device of claim 10 , further comprising:

a module configured to track whether each lexical pronunciation is either machine generated or human entered; and

a module configured to flag machine generated lexical pronunciations for human inspection.

17. The computing device of claim 10 , further comprising a module configured to add default pronunciations to the pronunciation dictionary based on TTS letter-to-sound rules.

18. A tangible computer-readable storage medium storing instructions for controlling a computing device for generating a database for a text-to-speech (TTS) voice, the instructions comprising:

matching via a processor every spoken word associated with a TTS voice database with a smallest set of possible pronunciations for each word, the smallest set being generated by:

automatically determining a dialect and linguistic context using linguistic rules;

empirically determining idiosyncratic speaker characteristics; and

determining a subject domain; and

dynamically generating a pronunciation dictionary on a word-by-word basis using the smallest set.

19. The computer-readable medium of claim 18 , wherein the dynamically generated pronunciation dictionary accounts for reading errors and idiosyncrasies of speakers.

20. The computer-readable medium of claim 19 , the instructions further comprising:

forcing an automatic speech recognition (ASR) module to choose from a subset of one or more variants of a word when more than one pronunciation variants exists for a given word.

Assignments (10)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038275/0238 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038275/0310 →