IP Library Granted Patent US 10,497,362
Granted Patent B2
US 10,497,362 · App. 15/905,728 · Granted Dec 3, 2019

System and method for outlier identification to remove poor alignments in speech synthesis

Inventors: E. Veera Raghavendra (Hyderabad, IN); Aravind Ganapathiraju (Hyderabad, IN)
G10L13/08G10L13/00G10L25/03G10L25/51G10L13/06G10L13/07G10L15/144
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,497,362
App. No.
15/905,728
Granted
Dec 3, 2019
Kind
B2
Abstract

A system and method are presented for outlier identification to remove poor alignments in speech synthesis. The quality of the output of a text-to-speech system directly depends on the accuracy of alignments of a speech utterance. The identification of mis-alignments and mis-pronunciations from automated alignments may be made based on fundamental frequency methods and group delay based outlier methods. The identification of these outliers allows for their removal, which improves the synthesis quality of the text-to-speech system.

Claims (59)

1. A method for identifying outlying results in audio files used for model training Hidden Markov Models, in a text-to-speech system, applying fundamental frequency, the method comprising the steps of:

a. extracting a plurality of values from the audio files of a text-to-speech system, wherein the values comprise fundamental frequencies;

b. generating alignments using the extracted plurality of values from the audio files in the text-to-speech system;

c. separating out instances of phonemes;

d. determining, for each separated instance, an average fundamental frequency value and an average duration value;

e. identifying an instance as an outlier, wherein an outlier is identified if:

i. the phoneme is a vowel;

ii. the average fundamental frequency of an instance is less than a predetermined value;

iii. the duration of the instance is greater than twice the average duration of a phoneme; and

iv. the duration of the instance is less than half of the average duration of a phoneme; and

f. determining a total number of the instances identified as outliers from step (e) for each sentence in the audio files;

g. discarding a sentence in the audio files in the text-to-speech system from training a Hidden Markov Model if the total number of the instances identified as outliers for the sentence is greater than a predetermined sentence outlier threshold.

2. The method of claim 1 , wherein the extracting is performed using a pitch tracking tool.

3. The method of claim 1 , wherein the generating is performed using a speech recognition system.

4. The method of claim 1 , wherein the alignments are generated at a phoneme level.

5. The method of claim 1 , wherein the threshold of outliers is

a predetermined value, and

the predetermined value of the threshold is empirically chosen.

6. The method of claim 1 , wherein the number of outliers of step (f) is empirically chosen.

7. The method of claim 6 , wherein the number of outliers is five.

8. A method for identifying outlying results in audio files used for model training Hidden Markov Models, in a text-to-speech system, applying group delay algorithms, the method comprising the steps of:

a. generating alignments of the audio files of the text-to-speech system at a phoneme level;

b. generating alignments of the audio files of the text-to-speech system at a syllable level;

c. adjusting the alignments at the syllable level using the group delay algorithms;

d. separating each syllable from the audio files of the text-to-speech system into a separate audio file;

e. generating, for each separate audio file of the text-to-speech system, phonemes of the separate audio files using phoneme boundaries for each syllable and an existing phoneme model;

f. determining a likelihood value of each generated phoneme, wherein if the likelihood value meets a criteria, identifying the generated phoneme as an outlier; and

g. determining a total number of outliers for each sentence in the audio files in the text-to-speech system;

h. discarding a sentence in the audio files in the text-to-speech system from training a Hidden Markov Model if the total number of the instances identified as outliers for the sentence is greater than a predetermined sentence outlier threshold.

9. The method of claim 8 , wherein the generating of step (a) is performed using a speech recognition system.

10. The method of claim 8 , wherein the generating of step (b) is performed using at least one of: a speech recognition system and a phoneme model.

11. The method of claim 8 , wherein the criteria comprises a small value.

12. The method of claim 8 , wherein the criteria comprises a failed alignment.

13. The method of claim 8 , wherein the number of outliers of step (g) comprises three.

14. A method for synthesizing speech in a text-to-speech system, wherein the system comprises at least a speech database, a database capable of storing Hidden Markov Models, and a synthesis filter, the method comprising the steps of:

a. identifying outlying results in audio files from the speech database and removing the outlying results before model training, comprising the steps of:

i. extracting a plurality of values from the audio files, wherein the values comprise fundamental frequencies,

ii. generating alignments using the extracted plurality of values from the audio files,

iii. separating out instances of phonemes,

iv. determining, for each separated instance, an average fundamental frequency value and an average duration value,

v. identifying an instance as an outlier, wherein an outlier is identified if:

(a) the phoneme is a vowel;

(b) the average fundamental frequency of an instance is less than a predetermined value;

(c) the duration of the instance is greater than twice the average duration of a phoneme; and

(d) the duration of the instance is less than half of the average duration of a phoneme, and

vi. determining a total number of the instances identified in step (a)(v) as outliers for each sentence in the audio files;

vii. discarding a sentence in the audio files from model training if the total number of the instances identified as outliers for the sentence is greater than a predetermined sentence outlier threshold,

b. converting a speech signal from the speech database into parameters and extracting the parameters from the speech signal;

c. training Hidden Markov Models using the extracted parameters from the speech signal and using the labels from the speech database to produce context dependent Hidden Markov Models;

d. storing the context dependent Hidden Markov Models in the database capable of storing Hidden Markov Models;

e. inputting text and analyzing the text, wherein said analyzing comprises extracting labels from the text;

f. utilizing said labels to generate parameters from the context dependent Hidden Markov Models;

g. generating an other signal from the parameters;

h. inputting the other signal and the parameters into the synthesis filter; and

i. producing synthesized speech as the other signal passes through the synthesis filter.

15. The method of claim 14 , wherein the parameters of step (b) comprise one or more of excitation and spectral.

16. The method of claim 14 , wherein the parameters of step (f) comprise one or more of excitation and spectral.

17. The method of claim 14 , wherein the extracting is performed using a pitch tracking tool.

18. The method of claim 14 , wherein the generating is performed using a speech recognition system.

Assignments (4)
CHANGE OF NAME Recorded Jun 6, 2024
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 067646/0448 →
SECURITY AGREEMENT Recorded Feb 12, 2020
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: BANK OF AMERICA, N.A.
Reel/Frame 051902/0850 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2019
From: RAGHAVENDRA, VEERA E.; GANAPATHIRAJU, ARAVIND
To: INTERACTIVE INTELLIGENCE GROUP, INC.
Reel/Frame 050221/0224 →
MERGER Recorded Aug 30, 2019
From: INTERACTIVE INTELLIGENCE GROUP, INC.
To: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
Reel/Frame 050221/0238 →
Continuity (2)
Continuation 14737080 · Jun 11, 2015
Related Publication 20180190265A1 · Jul 5, 2018