IP Library Granted Patent US 7,853,452
Granted Patent B2
US 7,853,452 · App. 12/327,579 · Granted Dec 14, 2010

Interactive debugging and tuning of methods for CTTS voice building

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,853,452
App. No.
12/327,579
Granted
Dec 14, 2010
Kind
B2
Abstract

A method, a system, and an apparatus for identifying and correcting sources of problems in synthesized speech which is generated using a concatenative text-to-speech (CTTS) technique. The method can include the step of displaying a waveform corresponding to synthesized speech generated from concatenated phonetic units. The synthesized speech can be generated from text input received from a user. The method further can include the step of displaying parameters corresponding to at least one of the phonetic units. The method can include the step of displaying the original recordings containing selected phonetic units. An editing input can be received from the user and the parameters can be adjusted in accordance with the editing input.

Claims (28)

1. A system for debugging and tuning synthesized audio, comprising:

means for receiving a user-supplied text with a visual user interface;

means for generating synthesized audio generated from concatenated phonetic units, the synthesized audio being a voice rendering of the user-supplied text;

means for displaying the waveform corresponding to synthesized audio generated from concatenated phonetic units;

means for displaying parameters corresponding to at least one of the phonetic units, the parameters including configuration parameters comprising at least one weight for adjusting at least one search cost function, the at least one weight comprising at least one of a pitch cost weight and a duration cost weight;

means for displaying an original recording containing a selected phonetic unit;

means for receiving an editing input from the user; and means for adjusting the parameters in accordance with the editing input by adjusting and storing in a text-to-speech engine configuration file at least one configuration parameter, wherein adjusting includes repositioning a phonetic alignment marker;

means for highlighting in the display of the original recording at least one user-selected phonetic unit;

means for correcting elements of a text-to-speech segment dataset of parameters corresponding to a segment of the synthesized audio identified as being problematic;

means for generating a new synthesized waveform corresponding to one or more adjusted parameters; and

wherein the system continues to regenerate new synthesized waveforms until a desired synthesized output is generated.

2. A machine-readable storage having stored thereon a computer program having a plurality of code sections, the code sections executable by a machine for causing the machine to perform the steps of:

(a) receiving a user-supplied text with a visual user interface;

(b) generating synthesized audio generated from concatenated phonetic units, the synthesized audio being a voice rendering of the user-supplied text;

(c) displaying a waveform corresponding to the synthesized audio generated from concatenated phonetic units;

(d) displaying parameters corresponding to at least one of the phonetic units, the parameters including configuration parameters comprising at least one weight for adjusting at least one search cost function, the at least one weight comprising at least one of a pitch cost weight and a duration cost weight;

(e) displaying an original recording containing a selected phonetic unit;

(f) receiving an editing input from the user;

(g) adjusting at least one configuration parameter in accordance with the editing input and storing the at least one configuration parameter in a text-to-speech engine configuration file, wherein adjusting includes repositioning a phonetic alignment marker;

(h) highlighting in the display of the original recording at least one user-selected phonetic unit;

(i) correcting elements of a text-to-speech segment dataset of parameters corresponding to a segment of the synthesized audio identified as being problematic;

(j) generating a new synthesized waveform corresponding to one or more adjusted parameters; and

(k) repeating steps (b)-(j) until a desired synthesized output is generated.

3. The machine-readable storage of claim 2 , wherein said displaying parameters step further comprises displaying the parameters responsive to a user selection of at least a portion of the waveform, the displayed parameters correlating to the selected portion of the waveform.

4. The machine-readable storage of claim 2 , wherein said displaying parameters step further comprises identifying a portion of the waveform responsive to a user selection of at least one of the parameters, the identified portion of the waveform correlating to the selected parameters.

5. The machine-readable storage of claim 2 , wherein said adjusting step comprises at least one action selected from the group consisting of deleting a pitch mark, inserting a pitch mark, and repositioning a pitch mark by deleting a phonetic unit label, adding a phonetic unit label, modifying the phonetic unit label, and repositioning the phonetic unit boundaries.

6. The machine-readable storage of claim 2 , wherein said displaying parameters step further comprises the step of displaying a waveform from the original recording along with the phonetic unit.

7. The machine-readable storage of claim 6 , wherein edits to the waveform adjust parameters in the segment dataset.

Assignments (9)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2009
From: GLEASON, PHILIP; SMITH, MARIA E.; VISWANATHAN, MAHESH; ZENG, JIE Z.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 022690/0603 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →