IP Library Granted Patent US 7,742,919
Granted Patent B1
US 7,742,919 · App. 11/235,857 · Granted Jun 22, 2010

System and method for repairing a TTS voice database

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,742,919
App. No.
11/235,857
Granted
Jun 22, 2010
Kind
B1
Abstract

The present invention provides various elements of a toolkit used for generating a TTS voice for use in a spoken dialog system. The embodiments in each case may be in the form of the system, a computer-readable medium or a method for generating the TTS voice. One embodiment of the invention relates to a method of correcting a database associated with the development of a text-to-speech (TTS) voice. The method comprises generating a pronunciation dictionary for use with a TTS voice, generating a TTS voice to a stage wherein it is prepared to be tested before being deployed, identifying mislabeled phonetic units associated with the TTS voice, for each identified mislabeled phonetic unit, linking to an entry within the pronunciation dictionary to correct the entry and deleting utterances and all associated data for unacceptable utterances.

Claims (40)

1. A computer implemented method of correcting a database associated with the development of a text-to-speech (TTS) voice, the method comprising:

generating via a processor a pronunciation dictionary for use with a TTS voice;

generating via the processor a TTS voice to a stage wherein it is prepared to be tested before being deployed;

receiving a single user input to identify all mislabeled phonetic units associated with the TTS voice at the stage wherein it is prepared to be tested before being deployed;

for each identified mislabeled phonetic unit, linking without additional user input to an entry within the pronunciation dictionary to correct the entry; and

deleting, without additional user input, from the pronunciation dictionary utterances and all associated data for unacceptable utterances.

2. The method of claim 1 , wherein the associated data is at least one of text, audio and labels.

3. The method of claim 1 , wherein deleting utterances and all associated data is performed automatically.

4. The method of claim 1 , wherein deleting utterances and all associated data occurs further for those utterances that cannot be successfully aligned by automatic speech recognition (ASR).

5. The method of claim 1 , further comprising:

correcting speaker dependent entries in the pronunciation database; and

rerunning ASR on all utterances containing the offending word.

6. The method of claim 5 , wherein a voice-builder module can automatically review only utterances that contain the offending word.

7. A computing device for correcting a database associated with the development of a text-to-speech (TTS) voice, the computing device comprising:

a processor;

a module configured to control the processor to generate a pronunciation dictionary for use with a TTS voice;

a module configured to control the processor to generate a TTS voice to a stage wherein it is prepared to be tested before being deployed;

a module configured to control the processor to receive a single user input to identify all mislabeled phonetic units associated with the TTS voice at the stage wherein it is prepared to be tested before being deployed;

a module configured to control the processor, for each identified mislabeled phonetic unit, to link to without additional user input an entry within the pronunciation dictionary to correct the entry; and

a module configured to control the processor to delete, without additional user input, from the pronunciation dictionary utterances and all associated data for unacceptable utterances.

8. The computing device of claim 7 , wherein the associated data is at least one of text, audio and labels.

9. The computing device of claim 7 , wherein deleting utterances and all associated data is performed automatically.

10. The computing device of claim 7 , wherein deleting utterances and all associated data occurs further for those utterances that cannot be successfully aligned by automatic speech recognition (ASR).

11. The computing device of claim 7 , further comprising:

a module configured to control the processor to correct speaker dependent entries in the pronunciation database; and

a module configured to control the processor to rerun ASR on all utterances containing the offending word.

12. The computing device of claim 11 , wherein a voice-builder module can automatically review only utterances that contain the offending word.

13. A non-transitory computer-readable storage medium storing instructions for controlling a computing device to correct a database associated with the development of a text-to-speech (TTS) voice, the instructions comprising:

generating via a processor a pronunciation dictionary for use with a TTS voice;

generating via a processor a TTS voice to a stage wherein it is prepared to be tested before being deployed;

receiving a single user input to identify mislabeled all phonetic units associated with the TTS voice at the stage wherein it is prepared to be tested before being deployed;

for each identified mislabeled phonetic unit, linking without additional user input to an entry within the pronunciation dictionary to correct the entry; and

deleting, without additional user input, from the pronunciation dictionary utterances and all associated data for unacceptable utterances.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the associated data is at least one of text, audio and labels.

15. The non-transitory computer-readable storage medium of claim 13 , wherein deleting utterances and all associated data is performed automatically.

16. The non-transitory computer-readable storage medium of claim 13 , wherein deleting utterances and all associated data occurs further for those utterances that cannot be successfully aligned by automatic speech recognition (ASR).

17. The non-transitory computer-readable storage medium of claim 13 , the instructions further comprising:

correcting speaker dependent entries in the pronunciation database; and

rerunning ASR on all utterances containing the offending word.

18. The non-transitory computer-readable storage medium of claim 17 , wherein a voice-builder module can automatically review only utterances that contain the offending word.

Assignments (10)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038275/0238 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038275/0310 →