IP Library Granted Patent US 7,454,348
Granted Patent B1
US 7,454,348 · App. 10/755,141 · Granted Nov 18, 2008

System and method for blending synthetic voices

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,454,348
App. No.
10/755,141
Granted
Nov 18, 2008
Kind
B1
Abstract

A system and method for generating a synthetic text-to-speech TTS voice are disclosed. A user is presented with at least one TTS voice and at least one voice characteristic. A new synthetic TTS voice is generated by blending a plurality of existing TTS voices according to the selected voice characteristics. The blending of voices involves interpolating segmented parameters of each TTS voice. Segmented parameters may be, for example, prosodic characteristics of the speech such as pitch, volume, phone durations, accents, stress, mis-pronunciations and emotion.

Claims (54)

1. A method of generating a synthetic voice comprising:

receiving a user selection of a first text-to-speech (TTS) voice and a second TTS voice from a plurality of TTS voices;

receiving at least one user-selected voice characteristic; and

generating a new TTS voice by blending the first TTS voice and the second TTS voice and according to the at least one user-selected voice characteristic.

2. The method of claim 1 , further comprising:

presenting the new TTS voice to the user for preview;

receiving user-selected adjustments; and

presenting a revised TTS voice to the user for preview according to the user-selected adjustments.

3. The method of claim 1 , wherein generating the new TTS voice further comprises interpolating between corresponding segment parameters of the first TTS voice and the second TTS voice.

4. The method of claim 3 , wherein the segment parameters relate to prosodic characteristics.

5. The method of claim 4 , wherein the prosodic characteristics are selected from a group comprising pitch contour, spectral envelope, volume contour and phone durations.

6. The method of claim 5 , wherein the prosodic characteristics are further selected from a group comprising syllable accent, language accent and emotion.

7. The method of claim 1 , wherein the user-selected voice characteristic relates to mis-pronunciations.

8. The method of claim 1 , wherein blending the first TTS voice and the second TTS voice further comprises extracting a prosodic characteristic from the LPC residual of the first TTS voice and the LPC residual of the second TTS voice and interpolating between the extracted prosodic characteristics.

9. The method of claim 8 , wherein the prosodic characteristics is pitch, wherein the interpolation of the extracted pitches from the first TTS voice and the second TTS voice generates a new blended pitch.

10. A method of generating a synthetic voice, the method comprising:

receiving a user selection of a TTS voice and a voice characteristic; and

presenting the user with a new TTS voice comprising the selected TTS voice blended with at least one other TTS voice to achieve the selected voice characteristics.

11. The method of claim 10 , further comprising:

presenting the new TTS voice to the user for preview;

receiving user-selected adjustments; and

presenting a revised TTS voice to the user for preview according to the user-selected adjustments.

12. The method of claim 10 , wherein generating the new TTS voice further comprises interpolating between corresponding segment parameters of the first TTS voice and the at least one other TTS voice.

13. The method of claim 11 , wherein the segment parameters relate to prosodic characteristics.

14. The method of claim 13 , wherein the prosodic characteristics are selected from a group comprising pitch contour, spectral envelope, volume contour and phone durations.

15. The method of claim 14 , wherein the prosodic characteristics are further selected from a group comprising: syllable accent, language accent and emotion.

16. The method of claim 10 , wherein the blended voice is generated by extracting a prosodic characteristic from the LPC residual of the first TTS voice and the LPC residual of the second TTS voice and interpolating between the extracted prosodic characteristics.

17. The method of claim 16 , wherein the prosodic characteristic is pitch and wherein the interpolation of the extracted pitches from the first TTS voice and the second TTS voice generates a new blended pitch.

18. The method of claim 10 , wherein the user-selected voice is blended with a plurality of other TTS voices to generate the new TTS voice.

19. The method of claim 10 , wherein the voice characteristic relates to mis-pronunciations.

20. A system for generating a synthetic voice, the system comprising:

a module for presenting a user with a plurality of TTS voices to select at least one voice characteristic;

a module for receiving a user-selected first TTS voice, a user-selected second TTS voice, and at least one user-selected voice characteristic; and

a module for generating a new TTS voice by blending the first TTS voice and the second TTS voice and according to the at least one user-selected voice characteristic.

21. The system of claim 20 , wherein the module that generates the new TTS voice further interpolates between corresponding segment parameters of the first TTS voice and the second TTS voice.

22. The system of claim 21 , wherein the segment parameters relate to prosodic characteristics.

23. The system of claim 22 , wherein the prosodic characteristics are selected from a group comprising pitch, contour, spectral envelope, volume contour and phone durations.

24. The system of claim 23 , wherein the prosodic characteristics are further selected from a group comprising: syllable accent, language accent and emotion.

25. The system of claim 20 , wherein blending the first TTS voice and the second TTS voice further comprises extracting a prosodic characteristic from the LPC residual of the first TTS voice and the LPC residual of the second TTS voice and interpolating between the extracted prosodic characteristics.

26. The system of claim 25 , wherein the prosodic characteristic is pitch, wherein the interpolation of the extracted pitches from the first TTS voice and the second TTS voice generates a new blended pitch.

27. A method of generating a text-to-speech (TTS) voice generated by blending at least two TTS voices, the method comprising:

establishing a voice profile for each of a plurality of TTS voices, each voice profile having speaker-specific parameters;

receiving a request for a new TTS voice from a user; and

generating the new TTS voice by blending speaker-specific parameters obtained from the voice profiles for at least two TTS voices.

28. The method of claim 27 , wherein the speaker-specific parameters comprise at least prosodic parameters associated with each TTS voice.

29. The method of claim 28 , wherein the speaker-specific parameters further comprise speaker-specific pronunciations.

30. The method of claim 27 , wherein the speaker-specific parameters are related to at least one of the group comprising: frame-based, phoneme-based, syllable-based and general characteristics.

31. A test-to-speech (TTS) voice generated from a method of blending at least two TTS voices, the method comprising:

establishing a voice profile for each of a plurality of TTS voices, each voice profile having speaker-specific parameters;

receiving a request for a blended TTS voice from a user; and

generating the blended TTS voice by blending speaker-specific parameters obtained from the voice profiles for at least two TTS voices.

32. The TTS voice of claim 31 , wherein the speaker-specific parameters comprise at least prosodic parameters associated with each TTS voice.

33. The TTS voice of claim 32 , wherein the speaker-specific parameters further comprise speaker-specific pronunciations.

34. The TTS voice of claim 33 , wherein the speaker-specific parameters are related to at least one of the group comprising: frame-based, phoneme-based, syllable-based and general characteristics.

Assignments (15)
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 036100/0925 Recorded Jul 1, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060559/0576 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0082 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0474 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded May 23, 2022
From: ORIX GROWTH CAPITAL, LLC
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 061749/0825 →
RELEASE OF SECURITY INTEREST Recorded May 18, 2020
From: BEARCUB ACQUISITIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 052693/0866 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0082 →
ASSIGNMENT OF IP SECURITY AGREEMENT Recorded Nov 17, 2017
From: ARES VENTURE FINANCE, L.P.
To: BEARCUB ACQUISITIONS LLC
Reel/Frame 044481/0034 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CHANGE PATENT 7146987 TO 7149687 PREVIOUSLY RECORDED ON REEL 036009 FRAME 0349. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 17, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 037134/0712 →
FIRST AMENDMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 13, 2015
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 036100/0925 →
SECURITY INTEREST Recorded Jun 23, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 036009/0349 →
SECURITY INTEREST Recorded Dec 19, 2014
From: INTERACTIONS LLC
To: ORIX VENTURES, LLC
Reel/Frame 034677/0768 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2014
From: AT&T ALEX HOLDINGS, LLC
To: INTERACTIONS LLC
Reel/Frame 034642/0640 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 034480/0960 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 034481/0031 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: AT&T ALEX HOLDINGS, LLC
Reel/Frame 034482/0414 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 8, 2004
From: KAPILOW, DAVID A.; ROSEN, KENNETH H.; SCHROETER, JUERGEN
To: AT&T CORPORATION
Reel/Frame 014941/0125 →