Method and apparatus for forced duration in neural speech synthesis
A system and method enable one to set a target duration of a desired synthesized utterance without removing or adding spoken content. Without changing the spoken text, the voice characteristics may be kept the same or substantially the same. Silence adjustment and interpolation may be used to alter the duration while preserving speech characteristics. Speech may be translated prior to a vocoder step, pursuant to which the translated speech is constrained by the original audio duration, while mimicking the speech characteristics of the original speech.
1. A system for constraining the duration of audio associated with text in a text-to-speech system, comprising:
ASR system producing a stream of text, and a time duration associated with the stream, as part of an end-to-end system;
a feature model configured to:
receive a text stream; and
produce a spectral feature output in frames associated with the text;
a duration constraint processor configured to:
receive the spectral feature output and the duration associated with the text stream; and
process the spectral feature output and the duration associated with the text stream to:
determine frames representing silence or an absence of text;
determine whether the stream is longer or shorter than the desired duration;
remove silence frames when required to reduce the duration of the spectral feature output;
add silence frames when required to increase the duration of the spectral feature output; and
perform interpolation on the spectral feature output after adjusting silence frames to make the duration of the spectral feature output match the required duration; and
a vocoder configured to:
receive the updated spectral feature output frames; and
produce synthesized audio.
2. The system of claim 1 , further comprising:
a machine translation engine, coupled to the ASR system and the feature model, configured to translate the text stream from the ASR system into a translated text stream,
wherein the feature model is further configured to:
receive the translated text stream, and
produce the spectral feature output corresponding to the translated text stream.
3. A method for constraining the duration of audio associated with text in a text-to-speech system, the method comprising:
producing a stream of text and a time duration associated with portions of the stream from an ASR system that is part of an end-to-end system;
receiving the text stream at a feature model;
generating a stream of spectral feature output in frames associated with the text;
determining frames representing silence or an absence of text;
determining whether the stream of spectral feature output is longer or shorter than the time duration associated with the text;
removing silence frames when required to reduce the duration of the spectral feature output;
adding silence frames when required to increase the duration of the spectral feature output;
performing interpolation on the spectral feature output after adjusting silence frames to make the duration of the spectral feature output match the required duration; and
synthesizing audio from the spectral feature output.
4. The method of claim 3 , wherein:
the silence frames are determined in segments of at least five frames; and
when silence frames are removed from the stream, the silence frames are removed evenly from the center of the segments.
5. The method of claim 4 , wherein, when silence frames are added to the stream, the silence frames are added evenly to the center of the segments.
6. The method of claim 3 , further comprising translating the text stream into another language using machine translation prior to receiving the text stream by the feature model.