IP Library Granted Patent US 7,702,510
Granted Patent B2
US 7,702,510 · App. 11/622,683 · Granted Apr 20, 2010

System and method for dynamically selecting among TTS systems

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,702,510
App. No.
11/622,683
Granted
Apr 20, 2010
Kind
B2
Abstract

Systems and methods for dynamically selecting among text-to-speech (TTS) systems. Exemplary embodiments of the systems and methods include identifying text for converting into a speech waveform, synthesizing said text by three TTS systems, generating a candidate waveform from each of the three systems, generating a score from each of the three systems, comparing each of the three scores, selecting a score based on a criteria and selecting one of the three waveforms based on the selected of the three scores.

Claims (44)

1. A method for dynamically selecting among text-to-speech (TTS) systems, the method comprising:

synthesizing a first section of text using a first TTS system employing a first algorithm to produce a first speech waveform having an associated first score;

synthesizing the first section of text using a second TTS system employing a second algorithm to produce a second speech waveform having an associated second score;

normalizing, with at least one processor configured to execute a normalizing function, the first score and the second score to produce a first normalized score and a second normalized score; and

selecting the first speech waveform or the second speech waveform for the first section of text based, at least in part, on a comparison of the first normalized score and the second normalized score.

2. The method as claimed in claim 1 , wherein the first score and the second score are cost function scores.

3. The method as claimed in claim 2 , wherein the speech waveform with the lowest cost function score is selected.

4. The method of claim 1 , wherein the first score and the second score are confidence scores.

5. The method of claim 1 , further comprising:

synthesizing a second section of text using the first TTS system to produce a third speech waveform having an associated third score;

synthesizing the second section of text using the second TTS system to produce a fourth speech waveform having an associated fourth score; and

selecting the third speech waveform or the fourth speech waveform for the second section of text based, at least in part, on a comparison of the third score and the fourth score;

wherein the speech waveform selected for the second section of text was synthesized using a different TTS system then the speech waveform selected for the first section of text.

6. The method of claim 5 , wherein the first section of text and second section of text are sub-sentence portions of text; and wherein the method further comprises:

concatenating the speech waveform selected for the first section of text with the speech waveform selected for the second section of text to form a concatenated speech waveform; and

outputting the concatenated speech waveform.

7. A system for dynamically selecting among text-to-speech (TTS) systems, comprising:

a plurality of TTS systems, each configured to receive a first section of text and to generate a first corresponding speech waveform having an associated first cost score;

at least one processor configured to normalize the associated first cost scores generated by the plurality of TTS systems to produce a plurality of normalized first cost scores; and

an output device configured to output one of said plurality of corresponding first speech waveforms having the lowest normalized first cost score from among the plurality of normalized first cost scores as speech output for said first section of text.

8. The system as claimed in claim 7 , wherein said plurality of TTS systems comprises a first TTS system employing a first TTS application and a second TTS system employing a second TTS application that is different than the first TTS application.

9. The system as claimed in claim 8 , wherein said first TTS application comprises a concatenative TTS engine and said second TTS application comprises a formant TTS engine.

10. The system of claim 7 , wherein the plurality of TTS systems are further configured to each receive a second section of text and to generate a corresponding second speech waveform having an associated second cost score; and

wherein the output device is further configured to output one of said plurality of corresponding second speech waveforms having the lowest associated second cost score from among the plurality of associated second cost scores as speech output for said second section of text;

wherein the speech waveform selected for the second section of text was synthesized using a different TTS system then the speech waveform selected for the first section of text.

11. The system of claim 10 , wherein the first section of text and second section of text are sub-sentence portions of text; and wherein the system further comprises:

a concatenation device configured to concatenate the speech waveform selected for the first section of text with the speech waveform selected for the second section of text to form a concatenated speech waveform; and

wherein the output device is further configured to output the concatenated speech waveform.

12. A computer-readable storage medium encoded with a plurality of instructions that, when executed by a computer, perform a method of dynamically selecting among text-to-speech (TTS) systems, the method, comprising:

synthesizing a first section of text using a first TTS system employing a first algorithm to produce a first speech waveform having an associated first score;

synthesizing the first section of text using a second TTS system employing a second algorithm to produce a second speech waveform having an associated second score;

normalizing the first score and the second score to produce a first normalized score and a second normalized score; and

selecting the first speech waveform or the second speech waveform based, at least in part, on a comparison of the first normalized score and the second normalized score.

13. The computer-readable storage medium of claim 12 , wherein the first score and the second score are cost function scores.

14. The computer-readable medium of claim 13 , wherein the speech waveform with the lowest cost function score is selected.

15. The computer-readable storage medium of claim 12 , wherein the first score and the second score are confidence scores.

16. The computer-readable storage medium of claim 12 , wherein the method further comprises:

synthesizing a second section of text using the first TTS system to produce a third speech waveform having an associated third score;

synthesizing the second section of text using the second TTS system to produce a fourth speech waveform having an associated fourth score; and

selecting the third speech waveform or the fourth speech waveform for the second section of text based, at least in part, on a comparison of the third score and the fourth score;

wherein the speech waveform selected for the second section of text was synthesized using a different TTS system then the speech waveform selected for the first section of text.

17. The computer-readable storage medium of claim 16 , wherein the first section of text and second section of text are sub-sentence portions of text; and wherein the method further comprises:

concatenating the speech waveform selected for the first section of text with the speech waveform selected for the second section of text to form a concatenated speech waveform; and

outputting the concatenated speech waveform.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →