Controllable diffusion-based speech generative model
Systems and techniques described herein relate to a diffusion-based model for generating converted speech from a source speech based on target speech. For example, a device may extract first prosody data from input data and may generate a content embedding based on the input data. The device may extract second prosody data from target speech, generate a speaker embedding from the target speech, and generate a prosody embedding from the second prosody data. The device may generate, based on the first prosody data and the prosody embedding, converted prosody data. The device may then generate a converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding.
1 . An apparatus to generate output speech from input data, comprising:
one or more memories configured to store the input data; and
one or more processors coupled to the one or more memories and configured to:
extract first prosody data from the input data;
generate a content embedding based on the input data;
extract second prosody data from target speech;
generate a speaker embedding from the target speech;
generate a prosody embedding from the second prosody data;
generate, based on the first prosody data and the prosody embedding, converted prosody data;
generate a converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding; and
generate converted speech based on converted spectrogram.
2 . The apparatus of claim 1 , wherein the input data comprises one or more of speech data or text data.
3 . The apparatus of claim 2 , wherein the input data comprises one of speech data and text data.
4 . The apparatus of claim 1 , wherein the first prosody data comprises one or more of a fundamental frequency, an energy value and a speed value.
5 . The apparatus of claim 1 , wherein the second prosody data comprises one or more of a fundamental frequency, an energy value and a speed value.
6 . The apparatus of claim 1 , wherein the one or more processors are configured to:
generate the converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding via a decoder comprising a diffusion decoder or a non-diffusion decoder.
7 . The apparatus of claim 1 , wherein the one or more processors are configured to:
generate, based on the converted prosody data, a predicted global speaking rate and via a rate control engine, a speaking rate for the converted spectrogram; and
generating, via a vocoder, the converted speech based on the input data.
8 . The apparatus of claim 7 , wherein the vocoder comprises a neural vocoder.
9 . The apparatus of claim 7 , wherein the rate control engine is configured to manipulate a speaking rate depending upon a predicted speed.
10 . The apparatus of claim 7 , wherein the one or more processors are configured to:
generate, based on the converted prosody data via the rate control engine, the speaking rate for the converted spectrogram independent of an automatic speech recognition model.
11 . The apparatus of claim 1 , wherein the one or more processors are configured to:
extract the first prosody data from the input data via a first prosody extractor engine;
generate the content embedding based on the input data via a content encoder;
extract the second prosody data from target speech via a second prosody extractor engine;
generate the speaker embedding from the target speech via a speaker encoder;
generate the prosody embedding from the second prosody data via a prosody encoder;
generate, based on the first prosody data and the prosody embedding, converted prosody data via a prosody conversion engine; and
generate the converted spectrogram based on the converted prosody data, the speaker embedding and the content embedding via a decoder.
12 . The apparatus of claim 11 , wherein the apparatus comprises the decoder, and wherein the decoder is configured to synthesize a speech spectrum conditioned on the content embedding, the speaker embedding, and the converted prosody data.
13 . The apparatus of claim 11 , the prosody encoder is configured to generate the prosody embedding at one or more of a frame-level or a sentence-level.
14 . The apparatus of claim 13 , wherein the apparatus comprises the prosody encoder, and wherein the prosody encoder is configured to generate the prosody embedding at the frame-level to enable frame-level intonation control.
15 . The apparatus of claim 1 , wherein the input data comprises speech data, the apparatus further comprising one or more microphones configured to capture the speech data.
16 . The apparatus of claim 1 , further comprising one or more speakers configured to output speech data comprising the converted prosody data.
17 . A method of generating output speech from input, the method comprising:
extracting first prosody data from input data;
generating a content embedding based on the input data;
extracting second prosody data from target speech;
generating a speaker embedding from the target speech;
generating a prosody embedding from the second prosody data;
generating, based on the first prosody data and the prosody embedding, converted prosody data;
generating a converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding; and
generating converted speech based on converted spectrogram.
18 . The method of claim 17 , wherein the input data comprises one or more of speech data or text data.
19 . The method of claim 18 , wherein the input data comprises one of speech data and text data.
20 . The method of claim 17 , wherein the first prosody data comprises one or more of a fundamental frequency, an energy value and a speed value.
21 . The method of claim 17 , wherein the second prosody data comprises one or more of a fundamental frequency, an energy value and a speed value.
22 . The method of claim 17 , further comprising:
generating the converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding via a decoder comprising a diffusion decoder or a non-diffusion decoder.
23 . The method of claim 17 , further comprising:
generating, based on the converted prosody data, a predicted global speaking rate and via a rate control engine, a speaking rate for the converted spectrogram; and
generating, via a vocoder, the converted speech based on the input data.
24 . The method of claim 23 , wherein the vocoder comprises a neural vocoder.
25 . The method of claim 23 , further comprising manipulating, using the rate control engine, a speaking rate depending upon a predicted speed.
26 . The method of claim 23 , further comprising:
generating, based on the converted prosody data via the rate control engine, the speaking rate for the converted spectrogram independent of an automatic speech recognition model.
27 . The method of claim 17 , further comprising:
extracting the first prosody data from the input data via a first prosody extractor engine;
generating the content embedding based on the input data via a content encoder;
extracting the second prosody data from target speech via a second prosody extractor engine;
generating the speaker embedding from the target speech via a speaker encoder;
generating the prosody embedding from the second prosody data via a prosody encoder;
generating, based on the first prosody data and the prosody embedding, converted prosody data via a prosody conversion engine; and
generating the converted spectrogram based on the converted prosody data, the speaker embedding and the content embedding via a decoder.
28 . The method of claim 27 , further comprising synthesizing, using a decoder, a speech spectrum conditioned on the content embedding, the speaker embedding, and the converted prosody data.
29 . The method of claim 27 , further comprising generating, using the prosody encoder, the prosody embedding at one or more of a frame-level or a sentence-level.
30 . The method of claim 29 , further comprising generating, using a prosody encoder, the prosody embedding at the frame-level to enable frame-level intonation control.