IP Library Granted Patent US 9,972,300
Granted Patent B2
US 9,972,300 · App. 14/737,080 · Granted May 15, 2018

System and method for outlier identification to remove poor alignments in speech synthesis

Inventors: E. Veera Raghavendra (Hyderabad, IN); Aravind Ganapathiraju (Hyderabad, IN)
Assignee: Genesys Telecommunications Laboratories, Inc.
G10L13/08G10L13/00G10L25/03G10L25/51G10L13/06G10L13/07G10L15/144
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,972,300
App. No.
14/737,080
Granted
May 15, 2018
Kind
B2
Abstract

A system and method are presented for outlier identification to remove poor alignments in speech synthesis. The quality of the output of a text-to-speech system directly depends on the accuracy of alignments of a speech utterance. The identification of mis-alignments and mis-pronunciations from automated alignments may be made based on fundamental frequency methods and group delay based outlier methods. The identification of these outliers allows for their removal, which improves the synthesis quality of the text-to-speech system.

Claims (36)

1. A method for generating synthesized speech using parametric models, the method comprising the steps of:

a. selecting sentences from a database of speech audio files, wherein the sentences comprise a plurality of phonemes;

b. identifying a total sum of instance outliers for each of the plurality of phonemes, wherein the instance outliers comprise fundamental frequency based outliers and group delay based outliers;

c. ignoring the sentences wherein the total sum of instance outliers exceeds a sentence outlier threshold and retaining sentences wherein the total sum of instance outliers meets the sentence outlier threshold;

d. using the retained sentences to generate trained Hidden Markov Models;

e. generating a plurality of context dependent Hidden Markov Models using the trained Hidden Markov Models, spectrum parameters, and excitation parameters, wherein the spectrum parameters and excitation parameters are extracted from the database of speech audio files using the trained Hidden Markov Models;

f. analyzing a selected text and generating text excitation parameters and text spectral parameters using the plurality of context dependent Hidden Markov Models;

g. generating a text excitation signal using the text excitation parameters; and

h. generating a synthesized speech waveform by passing the text excitation signal and text spectral parameters into a synthesis filter.

2. The method of claim 1 wherein the step of identifying the fundamental frequency based outliers further comprises:

a. performing signal analysis using a pitch tracking tool to extract the fundamental frequency of each of the sentences;

b. generating alignments using a speech recognition tool including a Hidden Markov Model Toolkit (HTK);

c. separating instances of the plurality of phonemes;

d. determining a first fundamental frequency and a duration from each of the separated instances of the plurality of phonemes; and

e. identifying instance outliers for each of the separated instances of the plurality of phonemes, wherein the instance outliers exceed an instance outlier threshold.

3. The method of claim 2 , wherein the instance outlier threshold is selected to identify phonemes presenting as vowels.

4. The method of claim 2 , wherein the instance outlier threshold is a predetermined fundamental frequency value, and an average of the first fundamental frequency of the separated instances of the plurality of phonemes is less than the predetermined fundamental frequency value.

5. The method of claim 4 wherein the predetermined fundamental frequency value is an empirically chosen value for each of the separated instances of the plurality of phonemes.

6. The method of claim 2 , wherein the instance outlier threshold is a duration when each of the separated instance presents at greater than twice an average duration of the phoneme.

7. The method of claim 2 , wherein the instance outlier threshold is a duration when each of the separated instance presents at less than half of the average duration of a phoneme.

8. The method of claim 1 wherein the step of identifying group delay based outliers further comprises:

a. generating syllable alignments for each of the plurality of phonemes using a speech recognition system and a phoneme model;

b. making adjustments to the syllable alignments using group delay algorithms;

c. splitting the syllable alignments and analyzing the split syllable assignments for pooling information;

d. generating phoneme boundaries for each of the split syllables using the phoneme model;

e. determining likelihood values for each of the generated phoneme boundaries, wherein the likelihood values comprise log-likelihood values;

f. determining whether generating the syllable alignment has failed or if the likelihood value is too small; and

g. identifying a sum of instance outliers for each of the generated phoneme boundaries.

9. The method of claim 8 , wherein the phoneme model comprises a previously trained acoustic model using training data.

10. The method of claim 1 , wherein the generating the plurality of context dependent Hidden Markov Models using the trained Hidden Markov Models further comprises extracting the spectrum parameters and the excitation parameters from the database of speech audio files and converting the spectrum parameters and the excitation parameters into a sequence of observed feature vectors.

11. The method of claim 1 , wherein the generating the text excitation parameters and the text spectral parameters further comprises the step of converting the selected text to a context-based label sequence.

12. The method of claim 1 , wherein the synthesis filter comprises filter parameters selected from the group consisting of Mel frequency cepstral coefficients, Mel frequency cepstral coefficients modeled by a statistical time series by using the plurality of context dependent Hidden Markov Models, and a first fundamental frequency.

13. The method of claim 1 , wherein the generating the synthesized speech waveform further comprises:

a. constructing a sentence Hidden Markov Model by concatenating a plurality of the context dependent Hidden Markov Models; and

b. determining a state duration for the sentence Hidden Markov Model, wherein the state duration is calculated to maximize an output probability of the state duration.

14. The method of claim 13 , wherein the output probability comprises a sequence of Mel frequency cepstral coefficients and log of first fundamental frequency values of separated instances of the plurality of phonemes.

Assignments (7)
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 04814/0387 Recorded Feb 5, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070115/0445 →
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 040815/0001 Recorded Feb 3, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070498/0001 →
CHANGE OF NAME Recorded Jun 6, 2024
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 067646/0452 →
SECURITY AGREEMENT Recorded Feb 22, 2019
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.; ECHOPASS CORPORATION; GREENEDEN U.S. HOLDINGS II, LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 048414/0387 →
MERGER Recorded Jul 1, 2018
From: INTERACTIVE INTELLIGENCE GROUP, INC.
To: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
Reel/Frame 046463/0839 →
SECURITY AGREEMENT Recorded Dec 5, 2016
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC., AS GRANTOR; ECHOPASS CORPORATION; INTERACTIVE INTELLIGENCE GROUP, INC.; BAY BRIDGE DECISION TECHNOLOGIES, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 040815/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2015
From: RAGHAVENDRA, E. VEERA; GANAPATHIRAJU, ARAVIND
To: INTERACTIVE INTELLIGENCE GROUP, INC.
Reel/Frame 035887/0546 →
Continuity (1)
Related Publication 20160365085A1 · Dec 15, 2016