IP Library Granted Patent US 10,991,360
Granted Patent B2
US 10,991,360 · App. 15/664,694 · Granted Apr 27, 2021

System and method for generating customized text-to-speech voices

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,991,360
App. No.
15/664,694
Granted
Apr 27, 2021
Kind
B2
Abstract

A system and method are disclosed for generating customized text-to-speech voices for a particular application. The method comprises generating a custom text-to-speech voice by selecting a voice for generating a custom text-to-speech voice associated with a domain, collecting text data associated with the domain from a pre-existing text data source and using the collected text data, generating an in-domain inventory of synthesis speech units by selecting speech units appropriate to the domain via a search of a pre-existing inventory of synthesis speech units, or by recording the minimal inventory for a selected level of synthesis quality. The text-to-speech custom voice for the domain is generated utilizing the in-domain inventory of synthesis speech units. Active learning techniques may also be employed to identify problem phrases wherein only a few minutes of recorded data is necessary to deliver a high quality TTS custom voice.

Claims (51)

1. A method comprising:

maintaining a pre-existing inventory of synthesis speech units;

collecting text data from a pre-existing text data source, to yield collected text data, wherein the collected text data is associated with a website;

converting the collected text data into linguistic tokens associated with the website, the linguistic tokens tagged with additional information selected from the group consisting of prosody, pronunciation, and speech acts and emotions;

generating an in-domain inventory of synthesis speech units as a first subset of the synthesis speech units of the pre-existing inventory of synthesis speech units including selecting synthesis speech units specific to the website from the pre-existing inventory of synthesis speech units, wherein the selecting occurs using the linguistic tokens associated with the website, to yield the first subset of synthesis speech units, wherein the synthesis speech units comprise one or more of phonemes, diphones, triphones and syllables;

generating, after generating the in-domain inventory of synthesis speech units, via a processor, a custom text-to-speech voice for use with the web site utilizing the in-domain inventory of synthesis speech units; and

providing, by the processor, an animation-based interaction with the website using the custom text-to-speech voice, the providing including generating a spoken response including generating a second subset of the synthesis speech units of the in-domain inventory of synthesis speech units based at least in part on a text generated by a language generator.

2. The method of claim 1 , further comprising:

caching the selected synthesis speech units to generate the in-domain inventory of synthesis speech units.

3. The method of claim 1 , further comprising:

determining whether the custom text-to-speech voice conforms to a selected level of synthesis quality.

4. The method of claim 3 , further comprising:

collecting additional text data associated with a domain if the custom text-to-speech voice does not conform to the selected level of synthesis quality.

5. The method of claim 4 , further comprising:

iteratively collecting the additional text data until the custom text-to-speech voice conforms to the selected level of synthesis quality.

6. The method of claim 1 , wherein the pre-existing text data source is one of a domain-related website, e-mail, and transcriptions of conversations.

7. The method of claim 1 , wherein the pre-existing text data source is a sector-related website distinct from the website.

8. The method of claim 7 , further comprising:

categorizing websites by sector to identify websites as pre-existing text data sources prior to collecting the text data.

9. The method of claim 1 , wherein collecting the text data from the pre-existing text data source further comprises mining specific words and phrases from the pre-existing text data source.

10. The method of claim 9 , wherein mining the specific words and phrases from the pre-existing text data source further comprises mining the specific words and phrases using an n-gram selection.

11. The method of claim 9 , wherein mining the specific words and phrases from the pre-existing text data source further comprises mining the specific words and phrases using a maximal mutual information approach.

12. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

maintaining a pre-existing inventory of synthesis speech units;

collecting text data from a pre-existing text data source, to yield collected text data, wherein the collected text data is associated with a web site;

converting the collected text data into linguistic tokens associated with the website, the linguistic tokens tagged with additional information selected from the group consisting of prosody, pronunciation, and speech acts and emotions;

generating an in-domain inventory of synthesis speech units as a first subset of the synthesis speech units of the pre-existing inventory of synthesis speech units including selecting synthesis speech units specific to the website from a pre-existing inventory of synthesis speech units, wherein the selecting occurs using the linguistic tokens associated with the website, to yield the first subset of synthesis speech units, wherein the synthesis speech units comprise one or more of phonemes, diphones, triphones and syllables;

generating, after generating the in-domain inventory of synthesis speech units, a custom text-to-speech voice for use with the website utilizing the in-domain inventory of synthesis speech units; and

providing an animation-based interaction with the website using the custom text-to-speech voice, the providing including generating a spoken response including generating a second subset of the synthesis speech units of the in-domain inventory of synthesis speech units based at least in part on a text generated by a language generator.

13. The system of claim 12 , wherein the computer-readable storage medium stores additional instructions stored which, when executed by the processor, cause the processor to perform operations further comprising:

caching the selected synthesis speech units to generate the in-domain inventory of synthesis speech units.

14. The system of claim 12 , wherein the computer-readable storage medium stores additional instructions stored which, when executed by the processor, cause the processor to perform operations further comprising:

determining whether the custom text-to-speech voice conforms to a selected level of synthesis quality.

15. The system of claim 14 , wherein the computer-readable storage medium stores additional instructions stored which, when executed by the processor, cause the processor to perform operations further comprising:

collecting additional text data associated with a domain if the custom text-to-speech voice does not conform to the selected level of synthesis quality.

16. The system of claim 15 , wherein the computer-readable storage medium stores additional instructions stored which, when executed by the processor, cause the processor to perform operations further comprising:

iteratively collecting the additional text data until the custom text-to-speech voice conforms to the selected level of synthesis quality.

17. The system of claim 12 , wherein the pre-existing text data source is one of a domain-related website, e-mail, and transcriptions of conversations.

18. The system of claim 12 , wherein the pre-existing text data source is a sector-related website distinct from the website.

19. The system of claim 18 , wherein the computer-readable storage medium stores additional instructions stored which, when executed by the processor, cause the processor to perform operations further comprising:

categorizing websites by sector to identify websites as pre-existing text data sources prior to collecting the text data.

20. A computer-readable storage device having instructions stored which, when executed by a processor, cause the processor to perform operations comprising:

maintaining a pre-existing inventory of synthesis speech units;

collecting text data from a pre-existing text data source, to yield collected text data, wherein the collected text data is associated with a website;

converting the collected text data into linguistic tokens associated with the website, the linguistic tokens tagged with additional information selected from the group consisting of prosody, pronunciation, and speech acts and emotions;

generating an in-domain inventory of synthesis speech units as a first subset of the synthesis speech units of the pre-existing inventory of synthesis speech units including selecting synthesis speech units specific to the website from the pre-existing inventory of synthesis speech units, wherein the selecting occurs using the linguistic tokens, to yield associated with the website, to yield the first subset of synthesis speech units, wherein the synthesis speech units comprise one or more of phonemes, diphones, triphones and syllables;

generating, after generating the in-domain inventory of synthesis speech units, a custom text-to-speech voice for use with the web site utilizing the in-domain inventory of synthesis speech units; and

providing an animation-based interaction with the website using the custom text-to-speech voice, the providing including generating a spoken response including generating a second subset of the synthesis speech units of the in-domain inventory of synthesis speech units based at least in part on a text generated by a language generator.

21. The method of claim 1 wherein generating the in-domain inventory of synthesis speech units further comprises receiving a recording of speech associated with at least some of the text data from the pre-existing text data source and processing the recording of speech to generate at least some synthesis speech units of the in-domain inventory of synthesis speech units.

Assignments (7)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →