IP Library Granted Patent US 12,670,895
Granted Patent B2
US 12,670,895 · App. 16/439,569 · Granted Jun 30, 2026

Invertible neural network to synthesize audio signals

Inventors: Ryan Prenger (Livermore, CA); Rafael Valle (Sunnyvale, CA); Bryan Catanzaro (Sunnyvale, CA)
Assignee: NVIDIA Corporation
G10L13/047G06N3/045G06N3/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,895
App. No.
16/439,569
Filed
Jun 12, 2019
Granted
Jun 30, 2026
Kind
B2
Art Unit
2658
USPC
706/15
Abstract

Systems and methods to help synthesize a second audio signal based, at least in part, on one or more neural networks trained using one or more characteristics of a first audio signal. Systems and methods to train one or more neural networks to synthesize a second audio signal based, at least in part, on one or more characteristics of a first audio signal.

Claims (71)

1 . One or more processors, comprising:

circuitry to:

train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are trained by:

converting a compact representation of a first audio signal;

generating one or more Gaussian values based at least in part on the converted compact representation and the first audio signal; and

training the one or more neural networks using the one or more Gaussian values; and

generate the one or more voice audio signals with the trained one or more neural networks.

2 . The one or more processors of claim 1 , wherein the compact representation is a mel-spectrogram.

3 . The one or more processors of claim 1 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.

4 . The one or more processors of claim 3 , wherein the one or more invertible layers include an audio transform that uses dilated convolutions to generate the one or more Gaussian values.

5 . The processor of claim 1 , wherein the one or more neural networks are to be trained in a first direction and to generate inferences in a second direction.

6 . A system, comprising:

one or more processors to use one or more circuits to:

train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are trained by:

converting a compact representation of a first audio signal;

generating one or more Gaussian values based, at least in part, on the converted compact representation and the first audio signal; and

training the one or more neural networks using the one or more Gaussian values; and

generate the one or more voice audio signals with the trained one or more neural networks.

7 . The system of claim 6 , wherein the one or more neural networks comprise one or more invertible layers.

8 . The system of claim 7 , wherein the one or more invertible layers include one or more audio transforms that use one or more dilated convolutions to generate the one or more Gaussian values.

9 . The system of claim 6 , wherein the compact representation is a mel-spectrogram.

10 . The system of claim 6 , wherein the first audio signal is human speech.

11 . The system of claim 6 , wherein the one or more neural networks are to be trained in a first direction and to generate inferences in a second direction.

12 . A speech synthesis system comprising:

one or more processors comprising one or more circuits to:

train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are trained by:

converting a compact representation of first audio signal;

generating one or more Gaussian values based, at least in part, on the converted compact representation and the first audio signal; and

training the one or more neural networks using the one or more Gaussian values; and

generate the one or more voice audio signals with the trained one or more neural networks.

13 . The speech synthesis system of claim 12 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.

14 . The speech synthesis system of claim 13 , wherein the one or more invertible layers comprise one or more invertible coupling layers comprising an audio transform function.

15 . The speech synthesis system of claim 12 , wherein the compact representation is a mel-spectrogram.

16 . The speech synthesis system of claim 12 , wherein the one or more neural networks generate Gaussian values in a first direction and synthesizes audio signals in a second direction.

17 . The speech synthesis system of claim 12 , wherein the speech synthesis system comprises a vehicle.

18 . One or more processors, comprising:

circuitry to:

train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are to be trained by at least:

converting a compact representation of a first audio signal;

generating one or more Gaussian values based, at least in part, on the converted compact representation and the first audio signal; and

training the one or more neural networks using the one or more Gaussian values; and

generate the one or more voice audio signals with the trained one or more neural networks.

19 . The one or more processors of claim 18 , wherein the compact representation is a mel-spectrogram.

20 . The one or more processors of claim 18 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.

21 . The one or more processors of claim 20 , wherein the one or more invertible layers include an audio transform that uses dilated convolutions to generate the one or more Gaussian values.

22 . The one or more processors of claim 18 , wherein the one or more neural networks are to be trained in a first direction and to inference in a second direction.

23 . A system, comprising:

one or more processors comprising one or more circuits to:

train one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein the one or more neural networks are to be trained by at least:

converting a compact representation of a first audio signal;

generating one or more Gaussian values based at least in part on the converted compact representation and the first audio signal; and

training the one or more neural networks using the one or more Gaussian values; and

generate the one or more voice audio signals with the trained one or more neural networks.

24 . The system of claim 23 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.

25 . The system of claim 24 , wherein the one or more invertible layers comprise one or more invertible coupling layers comprising an audio transform function.

26 . The system of claim 23 , wherein the one or more neural networks generate Gaussian values in a first direction and synthesize audio signals in a second direction.

27 . The system of claim 23 , wherein the one or more voice audio signals encodes synthesized human speech.

28 . The system of claim 23 , wherein a digital representation of the first audio signal is a digital recording of human speech.

29 . A computer-implemented method, comprising:

training one or more neural networks to infer one or more speech features of one or more voice audio signals to be generated based, at least in part, on one training that uses speech of the one or more voice audio signals to train the one or more neural networks to infer speech features, wherein training the one or more neural networks comprises:

converting a compact representation of a first audio signal;

generating one or more Gaussian values based, at least in part, on the converted compact representation and the first audio signal; and

training the one or more neural networks using the one or more Gaussian values; and

generating the one or more voice audio signals with the trained one or more neural networks.

30 . The method of claim 29 , wherein the compact representation is a mel-spectrogram.

31 . The method of claim 29 , wherein the one or more Gaussian values are generated using one or more invertible layers of the one or more neural networks.

32 . The method of claim 31 , wherein the one or more invertible layers include an audio transform that uses dilated convolutions to generate the one or more Gaussian values.

33 . The method of claim 29 , wherein the one or more neural networks are to be trained in a first direction and to inference in a second direction.

34 . The method of claim 29 , comprising:

training the one or more neural networks to generate one or more Gaussian values based, at least in part, on an audio signal; and

synthesizing an output voice audio signal based, at least in part, on different Gaussian values.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2019
From: PRENGER, RYAN; VALLE, RAFAEL; CATANZARO, BRYAN
To: NVIDIA CORPORATION
Reel/Frame 049756/0502 →
Continuity (1)
Related Publication 20200394994A1 · Dec 17, 2020
References Cited (39)
US 7254538B1 · Hermansky · 2007 [cited by examiner]
US 11631399B2 · Li · 2023 [cited by examiner]
US 20180018553A1 · Bach · 2018 [cited by examiner]
US 20190156837A1 · Park · 2019 [cited by examiner]
US 20190304480A1 · Narayanan · 2019 [cited by examiner]
US 20190355347A1 · Arik · 2019 [cited by examiner]
US 20200035247A1 · Boyadjiev · 2020 [cited by examiner]
US 20200066260A1 · Hayakawa · 2020 [cited by examiner]
US 20200074985A1 · Clark · 2020 [cited by examiner]
US 20200111496A1 · Itakura · 2020 [cited by examiner]
US 20200372361A1 · Ehteshami Bejnordi · 2020 [cited by examiner]
US 20200388267A1 · Bastyr · 2020 [cited by examiner]
US 20210375248A1 · Bonada · 2021 [cited by examiner]
US 20220076691A1 · Hiroya · 2022 [cited by examiner]
US 20220093104A1 · Sharifi · 2022 [cited by examiner]
US 20220208198A1 · Chang · 2022 [cited by examiner]
IEEE Computer Society, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” The Institute of Electrical and Electronics Engineers, Inc., Aug. 29, 2008, 70 pages. [cited by applicant]
Prenger et al., “WaveGlow: A Flow-based Generative Network for Speech Synthesis,” Oct. 31, 2018, retrieved Oct. 13, 2020, from https://arxiv.org/pdf/1811.00002.pdf, 5 pages. [cited by applicant]
Arik et al., “Deep Voice 2: Multi-Speaker Neural Textto-Speech,” Sep. 20, 2017, 15 pages. [cited by applicant]
Arik et al., “Deep Voice: Real-Time Neural Text-to-Speech,” Mar. 7, 2017, 17 pages. [cited by applicant]
Arik et al., “Fast Spectrogram Inversion Using Multi-Head Convolutional Neural Networks,” Nov. 6, 2018, 6 pages. [cited by applicant]
Dinh et al., “Density Estimation Using Real NVP,” Nov. 14, 2016, 30 pages. [cited by applicant]
Dinh et al., “NICE: Non-Linear Independent Components Estimation,” Dec. 19, 2014, 11 pages. [cited by applicant]
Griffin et al., “Signal Estimation from Modified Short-Time Fourier Transform,” IEEE Transactions on Acoustics, Speech, and Signal Processing, 32(2): Apr. 1984, 8 pages. [cited by applicant]
Ito et al., “The LJ Speech Dataset,” retrieved from https://keithito.com/LJ-Speech-Dataset/, 2017, 5 pages. [cited by applicant]
Jin et al., “FFTNET: A Real-Time Speaker-Dependent Neural Vocoder,” The 43rd IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 2018, 5 pages. [cited by applicant]
Kalchbrenner et al., “Efficient Neural Audio Synthesis,” Jun. 25, 2018, 10 pages. [cited by applicant]
Kingma et al. “Adam: A Method for Stochastic Optimization,” arXiv:1412.6980, dated Dec. 22, 2014, 9 pages. [cited by applicant]
Kingma et al., “Glow: Generative Flow with Invertible 1×1 Convolutions,” Jul. 10, 2018, 15 pages. [cited by applicant]
Kingma et al., “Improved Variational Inference with Inverse Autoregressive Flow,” Neural Information Processing Systems, 2016, 9 pages. [cited by applicant]
Oord et al., “Parallel WaveNet: Fast High-Fidelity Speech Synthesis,” Proceedings of the 35th International Conference on Machine Learning, 2018, 9 pages. [cited by applicant]
Oord et al., “WAVENET: A Generative Model for Raw Audio,” 2016, 15 pages. [cited by applicant]
Parmar et al., “Image Transformer,” ICML, 2018, 10 pages. [cited by applicant]
Ping et al., “ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech,” Jul. 30, 2018, 12 pages. [cited by applicant]
Rezende et al., “Variational Inference with Normalizing Flows,” Jun. 22, 2015, 10 pages. [cited by applicant]
Salimans et al., “Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks,” Advances in Neural Information Processing Systems, 2016, 9 pages. [cited by applicant]
Shen et al., “Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions,” Dec. 16, 2017, 5 pages. [cited by applicant]
Wang et al., “Tacotron: A Fully End-to-End Text-to-Speech Synthesis Model,” Apr. 6, 2017, 10 pages. [cited by applicant]
Yamamoto, “Wavenet Vocoder,” retrieved from https://zenodo.org/record/1472609#.YofMEHUpCUk, Oct. 27, 2018, 4 pages. [cited by applicant]