Invertible neural network to synthesize audio signals
Systems and methods to help synthesize a second audio signal based, at least in part, on one or more neural networks trained using one or more characteristics of a first audio signal. Systems and methods to train one or more neural networks to synthesize a second audio signal based, at least in part, on one or more characteristics of a first audio signal.
1 . One or more processors, comprising:
circuitry to:
train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are trained by:
converting a compact representation of a first audio signal;
generating one or more Gaussian values based at least in part on the converted compact representation and the first audio signal; and
training the one or more neural networks using the one or more Gaussian values; and
generate the one or more voice audio signals with the trained one or more neural networks.
2 . The one or more processors of claim 1 , wherein the compact representation is a mel-spectrogram.
3 . The one or more processors of claim 1 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.
4 . The one or more processors of claim 3 , wherein the one or more invertible layers include an audio transform that uses dilated convolutions to generate the one or more Gaussian values.
5 . The processor of claim 1 , wherein the one or more neural networks are to be trained in a first direction and to generate inferences in a second direction.
6 . A system, comprising:
one or more processors to use one or more circuits to:
train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are trained by:
converting a compact representation of a first audio signal;
generating one or more Gaussian values based, at least in part, on the converted compact representation and the first audio signal; and
training the one or more neural networks using the one or more Gaussian values; and
generate the one or more voice audio signals with the trained one or more neural networks.
7 . The system of claim 6 , wherein the one or more neural networks comprise one or more invertible layers.
8 . The system of claim 7 , wherein the one or more invertible layers include one or more audio transforms that use one or more dilated convolutions to generate the one or more Gaussian values.
9 . The system of claim 6 , wherein the compact representation is a mel-spectrogram.
10 . The system of claim 6 , wherein the first audio signal is human speech.
11 . The system of claim 6 , wherein the one or more neural networks are to be trained in a first direction and to generate inferences in a second direction.
12 . A speech synthesis system comprising:
one or more processors comprising one or more circuits to:
train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are trained by:
converting a compact representation of first audio signal;
generating one or more Gaussian values based, at least in part, on the converted compact representation and the first audio signal; and
training the one or more neural networks using the one or more Gaussian values; and
generate the one or more voice audio signals with the trained one or more neural networks.
13 . The speech synthesis system of claim 12 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.
14 . The speech synthesis system of claim 13 , wherein the one or more invertible layers comprise one or more invertible coupling layers comprising an audio transform function.
15 . The speech synthesis system of claim 12 , wherein the compact representation is a mel-spectrogram.
16 . The speech synthesis system of claim 12 , wherein the one or more neural networks generate Gaussian values in a first direction and synthesizes audio signals in a second direction.
17 . The speech synthesis system of claim 12 , wherein the speech synthesis system comprises a vehicle.
18 . One or more processors, comprising:
circuitry to:
train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are to be trained by at least:
converting a compact representation of a first audio signal;
generating one or more Gaussian values based, at least in part, on the converted compact representation and the first audio signal; and
training the one or more neural networks using the one or more Gaussian values; and
generate the one or more voice audio signals with the trained one or more neural networks.
19 . The one or more processors of claim 18 , wherein the compact representation is a mel-spectrogram.
20 . The one or more processors of claim 18 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.
21 . The one or more processors of claim 20 , wherein the one or more invertible layers include an audio transform that uses dilated convolutions to generate the one or more Gaussian values.
22 . The one or more processors of claim 18 , wherein the one or more neural networks are to be trained in a first direction and to inference in a second direction.
23 . A system, comprising:
one or more processors comprising one or more circuits to:
train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are to be trained by at least:
converting a compact representation of a first audio signal;
generating one or more Gaussian values based at least in part on the converted compact representation and the first audio signal; and
training the one or more neural networks using the one or more Gaussian values; and
generate the one or more voice audio signals with the trained one or more neural networks.
24 . The system of claim 23 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.
25 . The system of claim 24 , wherein the one or more invertible layers comprise one or more invertible coupling layers comprising an audio transform function.
26 . The system of claim 23 , wherein the one or more neural networks generate Gaussian values in a first direction and synthesize audio signals in a second direction.
27 . The system of claim 23 , wherein the one or more voice audio signals encodes synthesized human speech.
28 . The system of claim 23 , wherein a digital representation of the first audio signal is a digital recording of human speech.
29 . A computer-implemented method, comprising:
training one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on one training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein training the one or more neural networks comprises:
converting a compact representation of a first audio signal;
generating one or more Gaussian values based, at least in part, on the converted compact representation and the first audio signal; and
training the one or more neural networks using the one or more Gaussian values; and
generating the one or more voice audio signals with the trained one or more neural networks.
30 . The method of claim 29 , wherein the compact representation is a mel-spectrogram.
31 . The method of claim 29 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.
32 . The method of claim 31 , wherein the one or more invertible layers include an audio transform that uses dilated convolutions to generate the one or more Gaussian values.
33 . The method of claim 29 , wherein the one or more neural networks are to be trained in a first direction and to inference in a second direction.
34 . The method of claim 29 , comprising:
training the one or more neural networks to generate one or more Gaussian values based, at least in part, on an audio signal; and
synthesizing an output voice audio signal based, at least in part, on different Gaussian values.