IP Library Granted Patent US 9,721,558
Granted Patent B2
US 9,721,558 · App. 14/965,251 · Granted Aug 1, 2017

System and method for generating customized text-to-speech voices

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,721,558
App. No.
14/965,251
Granted
Aug 1, 2017
Kind
B2
Abstract

A system and method are disclosed for generating customized text-to-speech voices for a particular application. The method comprises generating a custom text-to-speech voice by selecting a voice for generating a custom text-to-speech voice associated with a domain, collecting text data associated with the domain from a pre-existing text data source and using the collected text data, generating an in-domain inventory of synthesis speech units by selecting speech units appropriate to the domain via a search of a pre-existing inventory of synthesis speech units, or by recording the minimal inventory for a selected level of synthesis quality. The text-to-speech custom voice for the domain is generated utilizing the in-domain inventory of synthesis speech units. Active learning techniques may also be employed to identify problem phrases wherein only a few minutes of recorded data is necessary to deliver a high quality TTS custom voice.

Claims (39)

1. A method comprising:

receiving, at a first time, a selection of an animated character to guide a user on a website;

collecting text data from a pre-existing text data source, to yield collected text data, wherein the collected text data is associated with a domain of the website, wherein the pre-existing text data source exists at the first time, and wherein no in-domain inventory of speech units exists at the first time;

selecting synthesis speech units specific to the domain from a pre-existing inventory of synthesis speech units existing at the first time, wherein the selecting occurs using the collected text data to yield selected synthesis speech units, wherein the synthesis speech units comprise one or more of phonemes, diphones, triphones and syllables;

caching the selected synthesis speech units specific to the domain to generate an in-domain inventory of synthesis speech units; and

generating, via a processor and at a second time which is later than the first time, a custom text-to-speech voice for use by the animated character when performing a specific task in the domain utilizing the in-domain inventory of synthesis speech units.

2. The method of claim 1 , further comprising determining whether the custom text-to-speech voice conforms to a selected level of synthesis quality.

3. The method of claim 2 , further comprising:

when the custom text-to-speech voice does not conform to the selected level of synthesis quality, collecting additional text data associated with the domain.

4. The method of claim 3 , further comprising iteratively collecting the additional text data until the custom text-to-speech voice conforms to the selected level of synthesis quality.

5. The method of claim 1 , wherein the pre-existing text data source is one of a domain-related website, e-mail, and transcriptions of conversations.

6. The method of claim 1 , wherein the pre-existing text data source is a sector-related website distinct from the website.

7. The method of claim 6 , further comprising categorizing websites by sector to identify websites as pre-existing text data sources prior to collecting the text data.

8. The method of claim 1 , wherein collecting the text data associated with the domain from the pre-existing text data source further comprises mining specific words and phrases from the pre-existing text data source.

9. The method of claim 8 , wherein mining the specific words and phrases from the pre-existing text data source further comprises mining the specific words and phrases using an n-gram selection.

10. The method of claim 8 , wherein mining the specific words and phrases from the pre-existing text data source further comprises mining the specific words and phrases using a maximal mutual information approach.

11. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

receiving, at a first time, a selection of an animated character to guide a user on a website;

collecting text data from a pre-existing text data source, to yield collected text data, wherein the collected text data is associated with a domain of the website, wherein the pre-existing text data source exists at the first time, and wherein no in-domain inventory of speech units exists at the first time;

selecting synthesis speech units specific to the domain from a pre-existing inventory of synthesis speech units existing at the first time, wherein the selecting occurs using the collected text data to yield selected synthesis speech units, wherein the synthesis speech units comprise one or more of phonemes, diphones, triphones and syllables;

caching the selected synthesis speech units specific to the domain to generate an in-domain inventory of synthesis speech units; and

generating, via a processor and at a second time which is later than the first time, a custom text-to-speech voice for use by the animated character when performing a specific task in the domain utilizing the in-domain inventory of synthesis speech units.

12. The system of claim 11 , further comprising determining whether the custom text-to-speech voice conforms to a selected level of synthesis quality.

13. The system of claim 12 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

when the custom text-to-speech voice does not conform to the selected level of synthesis quality, collecting additional text data associated with the domain.

14. The system of claim 13 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, cause the processor to perform operations comprising iteratively collecting the additional text data until the custom text-to-speech voice conforms to the selected level of synthesis quality.

15. The system of claim 11 , wherein the pre-existing text data source is one of a domain-related website, e-mail, and transcriptions of conversations.

16. The system of claim 11 , wherein the pre-existing text data source is a sector-related website distinct from the website.

17. The system of claim 16 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, cause the processor to perform operations comprising categorizing websites by sector to identify websites as pre-existing text data sources prior to collecting the text data.

18. The system of claim 11 , wherein collecting the text data associated with the domain from the pre-existing text data source further comprises mining specific words and phrases from the pre-existing text data source.

19. The system of claim 18 , wherein mining the specific words and phrases from the pre-existing text data source further comprises mining the specific words and phrases using an n-gram selection.

20. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

receiving, at a first time, a selection of an animated character to guide a user on a website;

collecting text data from a pre-existing text data source, to yield collected text data, wherein the collected text data is associated with a domain of the website, wherein the pre-existing text data source exists at the first time, and wherein no in-domain inventory of speech units exists at the first time;

selecting synthesis speech units specific to the domain from a pre-existing inventory of synthesis speech units existing at the first time, wherein the selecting occurs using the collected text data to yield selected synthesis speech units, wherein the synthesis speech units comprise one or more of phonemes, diphones, triphones and syllables;

caching the selected synthesis speech units specific to the domain to generate an in-domain inventory of synthesis speech units; and

generating, via a processor and at a second time which is later than the first time, a custom text-to-speech voice for use by the animated character when performing a specific task in the domain utilizing the in-domain inventory of synthesis speech units.

Assignments (11)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038529/0164 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038529/0240 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2016
From: BANGALORE, SRINIVAS; FENG, JUNLAN; RAHIM, MAZIN G.; SCHROETER, JUERGEN; SCHULZ, DAVID EUGENE; SYRDAL, ANN K.
To: AT&T CORP.
Reel/Frame 038294/0216 →