IP Library Granted Patent US 10,733,974
Granted Patent B2
US 10,733,974 · App. 15/874,612 · Granted Aug 4, 2020

System and method for synthesis of speech from provided text

Inventors: Yingyi Tan (Carmel, IN); Aravind Ganapathiraju (Hyderabad, IN); Felix Immanuel Wyss (Zionsville, IN)
G10L13/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,733,974
App. No.
15/874,612
Granted
Aug 4, 2020
Kind
B2
Abstract

A system and method are presented for the synthesis of speech from provided text. Particularly, the generation of parameters within the system is performed as a continuous approximation in order to mimic the natural flow of speech as opposed to a step-wise approximation of the feature stream. Provided text may be partitioned and parameters generated using a speech model. The generated parameters from the speech model may then be used in a post-processing step to obtain a new set of parameters for application in speech synthesis.

Claims (35)

1. A method for synthesizing speech from input text, the method comprising:

generating context labels from the input text, the context labels comprising one or more pause labels;

partitioning the input text into a plurality of linguistic segments in accordance with the one or more pause labels;

generating, for each linguistic segment, a time domain audio signal from the linguistic segment in accordance with a statistical parameter model;

generating, for each linguistic segment, a parameter trajectory from the time domain audio signal, the parameter trajectory comprising a plurality of frames for the linguistic segment, each frame comprising a vector of parameters;

smoothing a transition between a first frame and a second frame of the frames of the parameter trajectory; and

synthesizing speech from the parameter trajectory;

wherein the vector of parameters for each frame of the parameter trajectory comprises one or more frequency coefficients, spectral envelope values, delta coefficients, and delta-delta coefficients, and

wherein the smoothing the transition between the first frame and the second frame of the frames of the parameter trajectory comprises clamping at least one delta coefficient of the delta coefficients corresponding to the first frame and the second frame.

2. The method of claim 1 , wherein the context labels are generated based on linguistic analysis of the input text.

3. The method of claim 1 , wherein the frames of the parameter trajectory of the linguistic segment are grouped into a sequence of states, wherein the vectors of parameters for the frames are generated separately for each state of the sequence of states.

4. The method of claim 1 , wherein the generating the parameter trajectory comprises transforming the time domain audio signal to a spectral domain.

5. The method of claim 1 , wherein the statistical parameter model is trained by:

converting a speech corpus into a linguistic specification, the speech corpus covering sounds made in a language and the linguistic specification indexing the speech corpus to generate a speech waveform based on spectral speech parameters; and

generating the statistical parameter model based on the linguistic specification and a mean and covariance of a probability function fit by the spectral speech parameters.

6. A method for synthesizing speech from input text, the method comprising:

generating context labels from the input text, the context labels comprising one or more pause labels;

partitioning the input text into a plurality of linguistic segments in accordance with the one or more pause labels;

generating, for each linguistic segment, a time domain audio signal from the linguistic segment in accordance with a statistical parameter model;

generating, for each linguistic segment, a parameter trajectory from the time domain audio signal, the parameter trajectory comprising a plurality of frames for the linguistic segment, each frame comprising a vector of parameters;

smoothing a transition between a first frame and a second frame of the frames of the parameter trajectory; and

synthesizing speech from the parameter trajectory;

wherein the generating the parameter trajectory for a linguistic segment comprises generating a plurality of mel-cepstral coefficients by, for each frame of the parameter trajectory, where i is an index referring to a current frame:

setting a mel-cepstral coefficient of a first frame of the parameter trajectory to a mean value of a second frame of the parameter trajectory;

determining if the frame is voiced, wherein;

if the segment is unvoiced, setting the mel-cepstral coefficient of the current frame (mcep(i)) to (mcep(i−1)+mcep_mean(i))/2;

if the segment is voiced and is a first frame, then setting mcep(i)=(mcep(i−1)+mcep_mean(i))/2; and

if the segment is voiced and is not a first frame, then setting mcep(i)=(mcep(i−1)+mcep delta(i)+mcep_mean(i))/2;

determining if the linguistic segment has ended, wherein:

when the linguistic segment has ended, removing abrupt changes of the parameter trajectory and adjusting global variance; and

when the linguistic segment has not ended, incrementing the index i and repeating for the next frame of the parameter trajectory.

7. The method of claim 6 , wherein the statistical parameter model is trained by:

converting a speech corpus into a linguistic specification, the speech corpus covering sounds made in a language and the linguistic specification indexing the speech corpus to generate a speech waveform based on spectral speech parameters; and

generating the statistical parameter model based on the linguistic specification and a mean and covariance of a probability function fit by the spectral speech parameters.

8. The method of claim 6 , wherein the context labels are generated based on linguistic analysis of the input text.

Assignments (4)
CHANGE OF NAME Recorded Jun 6, 2024
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 067646/0448 →
SECURITY AGREEMENT Recorded Feb 12, 2020
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: BANK OF AMERICA, N.A.
Reel/Frame 051902/0850 →
MERGER Recorded Jul 1, 2018
From: INTERACTIVE INTELLIGENCE GROUP, INC.
To: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
Reel/Frame 046463/0839 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2018
From: TAN, YINGYI; GANAPATHIRAJU, ARAVIND; WYSS, FELIX IMMANUEL
To: INTERACTIVE INTELLIGENCE GROUP, INC.
Reel/Frame 045793/0301 →
Continuity (3)
Continuation 14596628 · Jan 14, 2015
Provisional Application 61927152 · Jan 14, 2014
Related Publication 20180144739A1 · May 24, 2018