IP Library Granted Patent US 8,438,032
Granted Patent B2
US 8,438,032 · App. 11/621,347 · Granted May 7, 2013

System for tuning synthesized speech

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,438,032
App. No.
11/621,347
Granted
May 7, 2013
Kind
B2
Abstract

An embodiment of the invention is a software tool used to convert text, speech synthesis markup language (SSML), and or extended SSML to synthesized audio. Provisions are provided to create, view, play, and edit the synthesized speech including editing pitch and duration targets, speaking type, paralinguistic events, and prosody. Prosody can be provided by way of a sample recording. Users can interact with the software tool by way of a graphical user interface (GUI). The software tool can produce synthesized audio file output in many file formats.

Claims (45)

1. A method of tuning synthesized speech, said method comprising:

synthesizing user supplied text to produce synthesized speech by a text-to-speech engine;

maintaining state information related to said synthesized speech;

receiving a user modification of duration cost factors associated with said synthesized speech to change the duration of said synthesized speech, including modifying a search of speech units when the text is re-synthesized to favor shorter speech units in response to user marking of any speech units in the synthesized speech as too long and modifying the search of speech units to favor longer speech units in response to user marking of any speech units in the synthesized speech as too short;

receiving a user modification of pitch cost factors associated with said synthesized speech to change the pitch of said synthesized speech;

receiving a user indication of segments of the user supplied text and/or the synthesized speech to skip during re-synthesis of said speech;

displaying a waveform associated with said synthesized speech and receiving user manipulations of the waveform; and

re-synthesizing said speech based on said user supplied text, said user modified duration cost factors, said user modified pitch cost factors, said user indicated segments to skip and said user manipulations of the waveform.

2. The method in accordance with claim 1 , further comprising:

highlighting, in response to a user input, a portion of a graphical representation of said synthesized speech.

3. The method in accordance with claim 2 , wherein highlighting further includes receiving a user selection of the highlighted portion to convert said synthesized speech to a SSML representation.

4. The method in accordance with claim 3 , further comprising:

adding a paralinguistic as SSML codes to said user supplied text.

5. The method in accordance with claim 4 , wherein said paralinguistic is at least one of the following:

i) a breath;

ii) a cough;

iii) a laugh;

iv) a sigh;

v) a throat clear; or

vi) a sniffle.

6. The method in accordance with claim 3 , further comprising:

adding a speaking style as SSML codes to said user supplied text.

7. The method in accordance with claim 6 , wherein said speaking style is apologetic.

8. The method in accordance with claim 6 , further comprising:

receiving a sample recording from said user to provide prosody.

9. The method in accordance with claim 1 , further comprising receiving a user indication of segments of the text that are to be used during re-synthesis of said speech.

10. A method of tuning synthesized speech, said method comprising:

synthesizing user supplied text to produce synthesized speech by a text-to-speech engine, said user supplied text including text, SSML or extended SSML;

displaying a waveform associated with said synthesized speech and receiving user manipulations of the waveform;

receiving a user modification of duration cost factors of said synthesized speech to change the duration of said synthesized speech;

receiving a user modification of pitch cost factors of said synthesized speech to change the pitch of said synthesized speech, including modifying a search of speech units when the text is re-synthesized to favor lower pitched speech units in response to user marking of any speech units in the synthesized speech as too high pitched and modifying the search of speech units to favor higher pitched speech units in response to user marking of any speech units in the synthesized speech as too low pitched;

receiving a user indication of segments of the user supplied text and/or the synthesized speech to skip during re-synthesis of said speech;

receiving a user indication of speech units to retain during re-synthesis of said speech; and

re-synthesizing said speech based on said user supplied text, said user modified duration cost factors, said user modified pitch cost factors, said user indicated segments to skip and said user manipulations of the waveform.

11. The method in accordance with claim 10 , further comprising:

highlighting, in response to a user input, a portion of a graphical representation of said synthesized speech.

12. The method in accordance with claim 11 , wherein highlighting further includes receiving a user selection of the highlighted portion to convert said synthesized speech to a SSML representation.

13. The method in accordance with claim 12 , further comprising:

adding a paralinguistic as SSML codes to said user supplied text.

14. The method in accordance with claim 13 , further comprising:

adding a speaking style as SSML codes to said user supplied text.

15. The method in accordance with claim 14 , further comprising:

receiving a sample recording from said user to provide prosody.

16. The method in accordance with claim 15 , wherein said waveform is a pitch contour of said synthesized speech.

17. The method in accordance with claim 10 , further comprising receiving a user indication of segments of the text, SSML or extended SSML that are to be used during re-synthesis of said speech.

Assignments (9)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2007
From: BAKIS, RAIMO; EIDE, ELLEN M.; PIERACCINI, ROBERTO; SMITH, MARIA E.; ZENG, JIE
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 018732/0893 →