IP Library Granted Patent US 8,666,746
Granted Patent B2
US 8,666,746 · App. 10/845,364 · Granted Mar 4, 2014

System and method for generating customized text-to-speech voices

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,666,746
App. No.
10/845,364
Granted
Mar 4, 2014
Kind
B2
Abstract

A system and method are disclosed for generating customized text-to-speech voices for a particular application. The method comprises generating a custom text-to-speech voice by selecting a voice for generating a custom text-to-speech voice associated with a domain, collecting text data associated with the domain from a pre-existing text data source and using the collected text data, generating an in-domain inventory of synthesis speech units by selecting speech units appropriate to the domain via a search of a pre-existing inventory of synthesis speech units, or by recording the minimal inventory for a selected level of synthesis quality. The text-to-speech custom voice for the domain is generated utilizing the in-domain inventory of synthesis speech units. Active learning techniques may also be employed to identify problem phrases wherein only a few minutes of recorded data is necessary to deliver a high quality TTS custom voice.

Claims (50)

1. A method comprising:

when a user is determined to be new, changing a front-end which converts text into linguistic tokens and tags the linguistic tokens with a prosody, a pronunciation, a speech act, and an emotion of the text, to yield a user-specific front-end;

in response to a user request from the user to generate a custom text-to-speech voice for a vehicle, the custom text-to-speech voice being associated with a domain, automatically:

collecting, via a processor, text data associated with the domain from a pre-existing text data source, to yield collected text data;

selecting, based on the user specific front end, synthesis speech units specific to the domain from a pre-existing inventory of synthesis speech units using the collected text data;

caching, via the processor, the synthesis speech units specific to the domain as an in-domain inventory of synthesis speech units; and

generating, via the processor, the custom text-to-speech voice for a specific task in the domain utilizing the in-domain inventory of synthesis speech units.

2. The method of claim 1 , further comprising determining whether the custom text-to-speech voice conforms to a selected level of synthesis quality.

3. The method of claim 2 , further comprising:

when the custom text-to-speech voice does not conform to the selected level of synthesis quality, collecting additional text data associated with the domain.

4. The method of claim 3 , further comprising iteratively collecting additional text data until the custom text-to-speech voice conforms to the selected level of synthesis quality.

5. The method of claim 1 , wherein the pre-existing text data source is one of a domain-related website, e-mail and transcriptions of conversations.

6. The method of claim 1 , wherein the pre-existing text data source is a sector-related website.

7. The method of claim 6 , further comprising categorizing websites by sector to identify websites as pre-existing text data sources prior to collecting the text data.

8. The method of claim 1 , wherein collecting text data associated with the domain from the pre-existing text data source further comprises mining specific words and phrases from the pre-existing text data source.

9. The method of claim 8 , wherein mining specific words and phrases from the pre-existing text data source further comprises mining specific words and phrases using an n-gram selection.

10. The method of claim 8 , wherein mining specific words and phrases from the pre-existing text data source further comprises mining specific words and phrases using a maximal mutual information approach.

11. The method of claim 1 , further comprising manually adding one of relevant words and relevant phrases to the collected text data for use in generating the in-domain inventory of synthesis speech units.

12. The method of claim 1 , further comprising applying active learning to identify one of problematic speech units and problematic phrases within the in-domain inventory of synthesis speech units.

13. The method of claim 12 , further comprising:

recording one of words and phrases according to the one of problematic speech units and problematic phrases; and

integrating the one of words and phrases into the in-domain inventory of synthesis speech units.

14. The method of claim 13 , further comprising:

determining whether the custom text-to-speech voice conforms to a selected synthesis quality.

15. The method of claim 14 , further comprising:

when the custom text-to-speech voice does not conform to the selected level of synthesis quality, collecting additional text data associated with the domain.

16. The method of claim 14 , wherein when the custom text-to-speech voice does not conform to the selected synthesis quality, recording one of additional words and additional phrases to increase the in-domain inventory of synthesis speech units.

17. The method of claim 12 , further comprising:

determining a minimal in-domain inventory for recording to meet a selected custom voice synthesis quality;

based on the minimal in-domain inventory, recording one of words and phrases according to the one of problematic speech units and problematic phrases; and

integrating the one of words and phrases into the in-domain task-independent inventory of synthesis speech units.

18. The method of claim 13 , further comprising:

determining whether the custom text-to-speech voice conforms to a custom voice synthesis quality; and

when the custom text-to-speech voice does not conform to the custom voice synthesis quality, recording one of additional words and additional phrases to increase the in-domain inventory of speech units.

19. A system comprising:

a processor; and

a computer readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

when a user is determined to be new, changing a front-end which converts text into linguistic tokens and tags the linguistic tokens with a prosody, a pronunciation, a speech act, and an emotion of the text, to yield a user-specific front-end;

in response to a user request from the user to generate a custom text-to-speech voice for a vehicle, the custom text-to-speech voice being associated with a domain, automatically:

collecting, via a processor, text data associated with the domain from a pre-existing text data source, to yield collected text data;

selecting, based on the user specific front end, synthesis speech units specific to the domain from a pre-existing inventory of synthesis speech units using the collected text data;

caching, via the processor, the synthesis speech units specific to the domain as an in-domain inventory of synthesis speech units; and

generating, via the processor, the custom text-to-speech voice for a specific task in the domain utilizing the in-domain inventory of synthesis speech units.

20. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

when a user is determined to be new, changing a front-end which converts text into linguistic tokens and tags the linguistic tokens with a prosody, a pronunciation, a speech act, and an emotion of the text, to yield a user-specific front-end;

in response to a user request from the user to generate a custom text-to-speech voice for a vehicle, the custom text-to-speech voice being associated with a domain, automatically:

collecting, via the computing device, text data associated with the domain from a pre-existing text data source, to yield collected text data;

selecting, based on the user specific front end, synthesis speech units specific to the domain from a pre-existing inventory of synthesis speech units using the collected text data;

caching, via the computing device, the synthesis speech units specific to the domain as an in-domain inventory of synthesis speech units; and

generating, via the computing device, the custom text-to-speech voice for a specific task in the domain utilizing the in-domain inventory of synthesis speech units.

Assignments (11)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2015
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 036737/0479 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2015
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 036737/0686 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2004
From: BANGALORE, SRINIVAS; FENG, JUNLAN; RAHIM, MAZIN G.; SCHROETER, JUERGEN; SCHULTZ, DAVID EUGENE; SYRDAL, ANN K.
To: AT&T CORP.
Reel/Frame 015339/0730 →