IP Library Granted Patent US 12670895
Granted Patent B2
US 12670895 · App. 16/439,569 · Granted Jun 30, 2026

Invertible neural network to synthesize audio signals

Inventors: Ryan Prenger (Livermore, CA); Rafael Valle (Sunnyvale, CA); Bryan Catanzaro (Sunnyvale, CA)
Assignee: NVIDIA Corporation
G10L13/047G06N3/045G06N3/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670895
App. No.
16/439,569
Granted
Jun 30, 2026
Kind
B2
Abstract

Systems and methods to help synthesize a second audio signal based, at least in part, on one or more neural networks trained using one or more characteristics of a first audio signal. Systems and methods to train one or more neural networks to synthesize a second audio signal based, at least in part, on one or more characteristics of a first audio signal.

Claims (71)

1 . One or more processors, comprising:

circuitry to:

train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are trained by:

converting a compact representation of a first audio signal;

generating one or more Gaussian values based at least in part on the converted compact representation and the first audio signal; and

training the one or more neural networks using the one or more Gaussian values; and

generate the one or more voice audio signals with the trained one or more neural networks.

2 . The one or more processors of claim 1 , wherein the compact representation is a mel-spectrogram.

3 . The one or more processors of claim 1 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.

4 . The one or more processors of claim 3 , wherein the one or more invertible layers include an audio transform that uses dilated convolutions to generate the one or more Gaussian values.

5 . The processor of claim 1 , wherein the one or more neural networks are to be trained in a first direction and to generate inferences in a second direction.

6 . A system, comprising:

one or more processors to use one or more circuits to:

train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are trained by:

converting a compact representation of a first audio signal;

generating one or more Gaussian values based, at least in part, on the converted compact representation and the first audio signal; and

training the one or more neural networks using the one or more Gaussian values; and

generate the one or more voice audio signals with the trained one or more neural networks.

7 . The system of claim 6 , wherein the one or more neural networks comprise one or more invertible layers.

8 . The system of claim 7 , wherein the one or more invertible layers include one or more audio transforms that use one or more dilated convolutions to generate the one or more Gaussian values.

9 . The system of claim 6 , wherein the compact representation is a mel-spectrogram.

10 . The system of claim 6 , wherein the first audio signal is human speech.

11 . The system of claim 6 , wherein the one or more neural networks are to be trained in a first direction and to generate inferences in a second direction.

12 . A speech synthesis system comprising:

one or more processors comprising one or more circuits to:

train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are trained by:

converting a compact representation of first audio signal;

generating one or more Gaussian values based, at least in part, on the converted compact representation and the first audio signal; and

training the one or more neural networks using the one or more Gaussian values; and

generate the one or more voice audio signals with the trained one or more neural networks.

13 . The speech synthesis system of claim 12 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.

14 . The speech synthesis system of claim 13 , wherein the one or more invertible layers comprise one or more invertible coupling layers comprising an audio transform function.

15 . The speech synthesis system of claim 12 , wherein the compact representation is a mel-spectrogram.

16 . The speech synthesis system of claim 12 , wherein the one or more neural networks generate Gaussian values in a first direction and synthesizes audio signals in a second direction.

17 . The speech synthesis system of claim 12 , wherein the speech synthesis system comprises a vehicle.

18 . One or more processors, comprising:

circuitry to:

train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are to be trained by at least:

converting a compact representation of a first audio signal;

generating one or more Gaussian values based, at least in part, on the converted compact representation and the first audio signal; and

training the one or more neural networks using the one or more Gaussian values; and

generate the one or more voice audio signals with the trained one or more neural networks.

19 . The one or more processors of claim 18 , wherein the compact representation is a mel-spectrogram.

20 . The one or more processors of claim 18 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.

21 . The one or more processors of claim 20 , wherein the one or more invertible layers include an audio transform that uses dilated convolutions to generate the one or more Gaussian values.

22 . The one or more processors of claim 18 , wherein the one or more neural networks are to be trained in a first direction and to inference in a second direction.

23 . A system, comprising:

one or more processors comprising one or more circuits to:

train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are to be trained by at least:

converting a compact representation of a first audio signal;

generating one or more Gaussian values based at least in part on the converted compact representation and the first audio signal; and

training the one or more neural networks using the one or more Gaussian values; and

generate the one or more voice audio signals with the trained one or more neural networks.

24 . The system of claim 23 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.

25 . The system of claim 24 , wherein the one or more invertible layers comprise one or more invertible coupling layers comprising an audio transform function.

26 . The system of claim 23 , wherein the one or more neural networks generate Gaussian values in a first direction and synthesize audio signals in a second direction.

27 . The system of claim 23 , wherein the one or more voice audio signals encodes synthesized human speech.

28 . The system of claim 23 , wherein a digital representation of the first audio signal is a digital recording of human speech.

29 . A computer-implemented method, comprising:

training one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on one training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein training the one or more neural networks comprises:

converting a compact representation of a first audio signal;

generating one or more Gaussian values based, at least in part, on the converted compact representation and the first audio signal; and

training the one or more neural networks using the one or more Gaussian values; and

generating the one or more voice audio signals with the trained one or more neural networks.

30 . The method of claim 29 , wherein the compact representation is a mel-spectrogram.

31 . The method of claim 29 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.

32 . The method of claim 31 , wherein the one or more invertible layers include an audio transform that uses dilated convolutions to generate the one or more Gaussian values.

33 . The method of claim 29 , wherein the one or more neural networks are to be trained in a first direction and to inference in a second direction.

34 . The method of claim 29 , comprising:

training the one or more neural networks to generate one or more Gaussian values based, at least in part, on an audio signal; and

synthesizing an output voice audio signal based, at least in part, on different Gaussian values.