Text-to-speech system with variable frame rate
A neural TTS system is trained to generate key acoustic frames at variable rates while omitting other frames. The frame skipping depends on the acoustic features to be generated for the input text. The TTS system can interpolate frames between the key frames at a target rate for a vocoder to synthesis audio samples.
1 . A computer-implemented method of speech synthesis, the method comprising:
receiving a sequence of symbols; and
synthesizing from the sequence of symbols, by a speech synthesis model, a plurality of key frames based on an average key frame rate input, wherein a key frame comprises at least one interpolation parameter indicating the ratio of the number of the plurality of key frames and one or more skipped frames is associated with the average key frame rate input.
2 . The computer-implemented method of claim 1 , wherein the at least one interpolation parameter indicates one or more skipped frames between the plurality of key frames.
3 . The computer-implemented method of claim 1 , wherein a frame rate of the plurality of key frames is variable and the at least one interpolation parameter indicates a length of time between the key frames.
4 . The computer-implemented method of claim 1 , wherein the at least one interpolation parameter comprises an indicated interpolation mode.
5 . The computer-implemented method of claim 1 , further comprising:
interpolating, by an interpolation model, one or more interpolated frames based on the plurality of key frames and the at least one interpolation parameter; and
generating, by a vocoder model, speech waveforms of the sequence of symbols based on the plurality of key frames and the one or more interpolated frames.
6 . The computer-implemented method of claim 1 , further comprising:
interpolating, by a vocoder model, one or more interpolated frames based on the plurality of key frames and the at least one interpolation parameter; and
generating, by the vocoder model, speech waveforms of the sequence of symbols based on the plurality of key frames and the one or more interpolated frames.
7 . A computer-implemented method of speech synthesis, the method comprising:
receiving, at a speech synthesis model, a sequence of symbols for speech synthesis;
synthesizing a plurality of key frames based on an average key frame rate input, wherein a key frame comprises at least one interpolation parameter indicating the ratio of the number of the plurality of key frames and the one or more skipped frames is associated with the average key frame rate input;
interpolating, by an interpolation model, one or more interpolated frames based on the plurality of key frames and the at least one interpolation parameter; and
generating, by a vocoder model, speech waveforms of the sequence of symbols based on the plurality of key frames and the one or more interpolated frames.
8 . The computer-implemented method of claim 7 , wherein the plurality of key frames have a variable frame rate.
9 . The computer-implemented method of claim 7 , wherein the at least one interpolation parameter indicates one or more skipped frames between the plurality of key frames.
10 . The computer-implemented method of claim 7 , wherein the at least one interpolation parameter indicates a length of time between the plurality of key frames.
11 . The computer-implemented method of claim 7 , wherein the at least one interpolation parameter comprises an indicated interpolation mode.
12 . A computer-implemented method of speech synthesis, the method comprising:
receiving, at a speech synthesis model, a sequence of symbols for speech synthesis;
synthesizing a plurality of key frames based on an average key frame rate input, wherein a key frame comprises at least one interpolation parameter indicating the ratio of the number of the plurality of key frames and the one or more skipped frames is associated with the average key frame rate input;
interpolating, by a vocoder model, one or more interpolated frames based on the plurality of key frames and the at least one interpolation parameter; and
generating, by the vocoder model, speech waveforms of the sequence of symbols based on the plurality of key frames and the one or more interpolated frames.
13 . The computer-implemented method of claim 12 , wherein the plurality of key frames have a variable frame rate.
14 . The computer-implemented method of claim 12 , wherein the at least one interpolation parameter indicates one or more skipped frames between the plurality of key frames.
15 . The computer-implemented method of claim 12 , wherein the at least one interpolation parameter indicates a length of time between the plurality of key frames.
16 . The computer-implemented method of claim 12 , wherein the at least one interpolation parameter comprises an indicated interpolation mode.