IP Library Granted Patent US 7,966,186
Granted Patent B2
US 7,966,186 · App. 12/264,622 · Granted Jun 21, 2011

System and method for blending synthetic voices

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,966,186
App. No.
12/264,622
Granted
Jun 21, 2011
Kind
B2
Abstract

A system and method for generating a synthetic text-to-speech TTS voice are disclosed. A user is presented with at least one TTS voice and at least one voice characteristic. A new synthetic TTS voice is generated by blending a plurality of existing TTS voices according to the selected voice characteristics. The blending of voices involves interpolating segmented parameters of each TTS voice. Segmented parameters may be, for example, prosodic characteristics of the speech such as pitch, volume, phone durations, accents, stress, mis-pronunciations and emotion.

Claims (42)

1. A tangible computer-readable medium storing instructions for controlling a computing device to generate a synthetic voice, the instructions comprising:

receiving a user selection of a first text-to-speech voice and a selected voice characteristic for modifying the first text-to-speech voice;

selecting the first text-to-speech voice from a plurality of text-to-speech voices;

selecting a second text-to-speech voice exhibiting the selected voice characteristic; and

presenting the user with a new text-to-speech voice comprising the first text-to-speech voice modified with at least the selected voice characteristic from the second text-to-speech voice.

2. The tangible computer-readable medium of claim 1 , the instructions further comprising:

presenting the new text-to-speech voice to the user for preview;

receiving user-selected adjustments; and

presenting a revised text-to-speech voice to the user for preview according to the user-selected adjustments.

3. The tangible computer-readable medium of claim 2 , wherein the segment parameters relate to prosodic characteristics.

4. The tangible computer-readable medium of claim 3 , wherein the prosodic characteristics are selected from a group comprising pitch contour, spectral envelope, volume contour and phone durations.

5. The tangible computer-readable medium of claim 4 , wherein the prosodic characteristics are further selected from a group comprising: syllable accent, language accent and emotion.

6. The tangible computer-readable medium of claim 1 , wherein generating the new text-to-speech voice further comprises interpolating between corresponding segment parameters of the first text-to-speech voice and the second text-to-speech voice.

7. The tangible computer-readable medium of claim 1 , wherein the new text-to-speech voice is generated by extracting a prosodic characteristic from a Linear-Predictive Coding residual of the first text-to-speech voice and the Linear-Predictive Coding residual of the second text-to-speech voice and interpolating between the extracted prosodic characteristics.

8. The tangible computer-readable medium of claim 7 , wherein the prosodic characteristic is pitch and wherein the interpolation of the extracted pitches from the first text-to-speech voice and the second text-to-speech voice generates a new blended pitch.

9. The tangible computer-readable medium of claim 1 , wherein the first text-to-speech voice is blended with a plurality of other text-to-speech voices to generate the new text-to-speech voice.

10. The tangible computer-readable medium of claim 1 , wherein the voice characteristic relates to mis-pronunciations.

11. A method of generating a synthetic voice, the method comprising:

receiving a user selection of a first text-to-speech voice and a selected voice characteristic for modifying the first text-to-speech voice;

selecting the first text-to-speech voice from a plurality of text-to-speech voices;

selecting a second text-to-speech voice exhibiting the selected voice characteristic; and

presenting the user with a new text-to-speech voice comprising the first text-to-speech voice modified with at least the selected voice characteristic from the second text-to-speech voice.

12. The method of claim 11 , wherein the first text-to-speech voice exhibiting the selected voice characteristic is generated by blending the first text-to-speech voice with the second text-to-speech voice.

13. The method of claim 12 , wherein the second text-to-speech voice includes the selected voice characteristic.

14. The method of claim 13 , wherein the new text-to-speech voice is generated to exhibit the selected voice characteristic by blending the first text-to-speech voice with at least the second text-to-speech voice.

15. The method of claim 11 , further comprising:

presenting the new text-to-speech voice to the user for preview;

receiving user-selected adjustments associated with the selected voice characteristic; and

presenting a revised text-to-speech voice for the user for preview according to the user selected adjustments to the selected voice characteristic.

16. The method of claim 11 , wherein the voice characteristic relates to mispronunciations.

17. A system for generating a synthetic voice, the system comprising:

a first module configured to control a processor to receive a user selection of a first text-to-speech voice and a selected voice characteristic for modifying the first text-to-speech voice;

a second module configured to control the processor to select the first text-to-speech voice from a plurality of text-to-speech voices;

a third module for configured to control the processor to select a second text-to-speech voice exhibiting the selected voice characteristic;

a fourth module configured to control the processor to present the user with a new text-to-speech comprising the first text-to-speech voice modified with the selected voice characteristic from the second text-to-speech voice.

18. The system of claim 17 , the system further comprising:

a fifth module configured to control the processor to present the new text-to-speech voice to the user for preview;

a sixth module configured to control the processor to receive user selected adjustments associated with a selected voice characteristic; and

a seventh module configured to control the processor to present a second new text-to-speech voice to the user for preview according to the user-selected adjustments of the selected voice characteristic.

19. The system of claim 18 , wherein each voice of the plurality of text-to-speech voices has speaker-specific parameters.

20. The system of claim 19 , wherein the speaker-specific parameters comprise at least prosodic parameters associated with each text-to-speech voice.

21. The system of claim 20 , wherein the speaker-specific parameters further comprise speaker-specific pronunciations.

Assignments (18)
RELEASE OF SECURITY INTEREST Recorded Sep 4, 2025
From: RUNWAY GROWTH FINANCE CORP., AS AGENT
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 072802/0931 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER PREVIOUSLY RECORDED AT REEL: 060445 FRAME: 0733. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 1, 2023
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 062919/0063 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 036100/0925 Recorded Jul 1, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060559/0576 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0082 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0474 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 27, 2022
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 060445/0733 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded May 23, 2022
From: ORIX GROWTH CAPITAL, LLC
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 061749/0825 →
RELEASE OF SECURITY INTEREST Recorded May 18, 2020
From: BEARCUB ACQUISITIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 052693/0866 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0082 →
ASSIGNMENT OF IP SECURITY AGREEMENT Recorded Nov 17, 2017
From: ARES VENTURE FINANCE, L.P.
To: BEARCUB ACQUISITIONS LLC
Reel/Frame 044481/0034 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CHANGE PATENT 7146987 TO 7149687 PREVIOUSLY RECORDED ON REEL 036009 FRAME 0349. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 17, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 037134/0712 →
FIRST AMENDMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 13, 2015
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 036100/0925 →
SECURITY INTEREST Recorded Jun 23, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 036009/0349 →
SECURITY INTEREST Recorded Dec 19, 2014
From: INTERACTIONS LLC
To: ORIX VENTURES, LLC
Reel/Frame 034677/0768 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2014
From: AT&T ALEX HOLDINGS, LLC
To: INTERACTIONS LLC
Reel/Frame 034642/0640 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: KAPILOW, DAVID A.; ROSEN, KENNETH H.; SCHROETER, JUERGEN
To: AT&T CORP.
Reel/Frame 034480/0900 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: AT&T ALEX HOLDINGS, LLC
Reel/Frame 034482/0414 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 034481/0031 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 034480/0960 →