IP Library › Granted Patent US 12,308,039
Granted Patent B2
US 12,308,039 · App. 17/687,266 · Granted May 20, 2025

Multi-band synchronized neural vocoder

Inventors: Chengzhu Yu (Bellevue, WA); Meng Yu (Bellevue, WA); Heng Lu (Sammamish, WA); Dong Yu (Bothell, WA)
Assignee: TENCENT AMERICA LLC
G10L19/16G06N3/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,308,039
App. No.
17/687,266
Granted
May 20, 2025
Kind
B2
Abstract

An apparatus and a method include receiving an input audio signal to be processed by a multi-band synchronized neural vocoder. The input audio signal is separated into a plurality of frequency bands. A plurality of audio signals corresponding to the plurality of frequency bands is obtained. Each of the audio signals is downsampled, and processed by the multi-band synchronized neural vocoder. An audio output signal is generated.

Claims (50)

1. A method performed by a multi-band synchronized neural vocoder, comprising:

receiving an input audio signal to be processed by the multi-band synchronized neural vocoder;

separating, by the multi-band synchronized neural vocoder, the input audio signal into a plurality of frequency bands;

obtaining, by the multi-band synchronized neural vocoder, a plurality of audio signals that corresponds to the plurality of frequency bands, based on separating the input audio signal into the plurality of frequency bands;

downsampling, by the multi-band synchronized neural vocoder, each of the plurality of audio signals, based on obtaining the plurality of audio signals;

processing, by the multi-band synchronized neural vocoder, the downsampled audio signals; and

generating, by the multi-band synchronized neural vocoder, an audio output signal based on processing the downsampled audio signals,

wherein, in the multi-band synchronized neural vocoder, each of the frequency bands has its own fully connected layer and a corresponding softmax layer, and

wherein weight parameters of the multi-band synchronized neural vocoder are shared across the plurality of frequency bands except for final fully connected layers and softmax layers for each of the frequency bands.

2. The method of claim 1 , wherein the downsampled audio signals of each of the plurality of frequency bands are processed simultaneously.

3. The method of claim 1 , wherein the downsampled audio signals of each of the plurality of frequency bands are processed using a single processing unit.

4. The method of claim 1 , wherein the neural vocoder is a WaveNet vocoder.

5. The method of claim 1 , wherein the neural vocoder is a WaveRNN vocoder.

6. The method of claim 1 , wherein the neural vocoder is an LPCNet vocoder.

7. The method of claim 1 , further comprising:

upsampling each of the processed audio signals; and

generating the audio output signal based on upsampling each of the processed audio signals.

8. A multi-band synchronized neural vocoder device, comprising:

at least one memory configured to store program code;

at least one processor configured to read the program code and operate as instructed by the program code, the program code including:

receiving code configured to cause that least one processor to receive an input audio signal to be processed by the multi-band synchronized neural vocoder;

separating code configured to cause the at least one processor to separate the input audio signal into a plurality of frequency bands;

obtaining code configured to cause the at least one processor to obtain a plurality of audio signals that corresponds to the plurality of frequency bands, based on separating the input audio signal into the plurality of frequency bands;

downsampling code configured to cause the at least one processor to downsample each of the plurality of audio signals, based on obtaining the plurality of audio signals;

processing code configured to cause the at least one processor to process the downsampled audio signals; and

generating code configured to cause the at least one processor to generate an audio output signal based on processing the downsampled audio signals,

wherein, in the multi-band synchronized neural vocoder, each of the frequency bands has its own fully connected layer and a corresponding softmax layer, and

wherein weight parameters of the multi-band synchronized neural vocoder are shared across the plurality of frequency bands except for final fully connected layers and softmax layers for each of the frequency bands.

9. The device of claim 8 , wherein the downsampled audio signals of each of the plurality of frequency bands are processed simultaneously.

10. The device of claim 8 , wherein the downsampled audio signals of each of the plurality of frequency bands are processed using a single processing unit.

11. The device of claim 8 , wherein the neural vocoder is a WaveNet vocoder.

12. The device of claim 8 , wherein the neural vocoder is a WaveRNN vocoder.

13. The device of claim 8 , wherein the neural vocoder is an LPCNet vocoder.

14. The device of claim 8 , further comprising:

upsampling code configured to cause the at least one processor to upsample each of the processed audio signals; and

wherein generating code is further configured to cause the at least one processor to generate the audio output signal based on upsampling each of the processed audio signals.

15. A non-transitory computer-readable medium storing instructions, the instructions comprising: one or more instructions that, when executed by one or more processors of a multi-band synchronized neural vocoder device, cause the one or more processors to:

receive an input audio signal to be processed by the multi-band synchronized neural vocoder device;

separate the input audio signal into a plurality of frequency bands;

obtain a plurality of audio signals that corresponds to the plurality of frequency bands, based on separating the input audio signal into the plurality of frequency bands;

downsample each of the plurality of audio signals, based on obtaining the plurality of audio signals;

process the downsampled audio signals; and

generate an audio output signal based on processing the downsampled audio signals,

wherein, in the multi-band synchronized neural vocoder, each of the frequency bands has its own fully connected layer and a corresponding softmax layer, and

wherein weight parameters of the multi-band synchronized neural vocoder are shared across the plurality of frequency bands except for final fully connected layers and softmax layers for each of the frequency bands.

16. The non-transitory computer-readable medium of claim 15 , wherein the downsampled audio signals of each of the plurality of frequency bands are processed simultaneously.

17. The non-transitory computer-readable medium of claim 15 , wherein the downsampled audio signals of each of the plurality of frequency bands are processed using a single processing unit.

18. The non-transitory computer-readable medium of claim 15 , wherein the neural vocoder is a WaveNet vocoder.

19. The non-transitory computer-readable medium of claim 15 , wherein the neural vocoder is a WaveRNN vocoder.

20. The non-transitory computer-readable medium of claim 15 , wherein the neural vocoder is an LPCNet vocoder.

Continuity (2)
Continuation 16576943 · Sep 20, 2019
Related Publication 20220189495A1 · Jun 16, 2022
References Cited (33)
US 5425130A · Morgan · 1995 [cited by applicant]
US 5715365A · Griffin et al. · 1998 [cited by applicant]
US 5809455A · Nishiguchi et al. · 1998 [cited by applicant]
US 6041297A · Goldberg · 2000 [cited by applicant]
US 6233550B1 · Gersho et al. · 2001 [cited by applicant]
US 6475245B2 · Gersho et al. · 2002 [cited by applicant]
US 8078474B2 · Vos et al. · 2011 [cited by applicant]
US 8566259B2 · Chong et al. · 2013 [cited by applicant]
US 9124981B2 · Grokop · 2015 [cited by applicant]
US 10529349B2 · Le Roux et al. · 2020 [cited by applicant]
US 20010023396A1 · Gersho et al. · 2001 [cited by applicant]
US 20110066578A1 · Chong et al. · 2011 [cited by applicant]
US 20140133663A1 · Grokop · 2014 [cited by applicant]
US 20140195227A1 · Rudzicz et al. · 2014 [cited by applicant]
US 20140201126A1 · Zadeh et al. · 2014 [cited by applicant]
US 20190066657A1 · Okamoto et al. · 2019 [cited by applicant]
US 20190122651A1 · Arik et al. · 2019 [cited by applicant]
US 20190318754A1 · Le Roux et al. · 2019 [cited by applicant]
Rabiee, Azam, et al. “A Fully Time-domain Neural Model for Subband-based Speech Synthesizer.” arXiv preprint arXiv:1810.05319 (2018). (Year: 2018). [cited by examiner]
Okamoto, Takuma, et al. “Improving FFTNet vocoder with noise shaping and subband approaches.” 2018 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2018. (Year: 2018). [cited by examiner]
Written Opinion in International Application No. PCT/US2020/045911, issued on Oct. 22, 2020. [cited by applicant]
International Search Report in International Application No. PCT/US2020/045911, issued on Oct. 22, 2020. [cited by applicant]
Arik et at. “Deep Voice: Real-time Neural Text-to-Speech,” arXiv:1702.07825v2 (cs CL] 7 Mar. 1-20, 2017, [retrieved on Oct. 11, 2020]. Retrieved from the Internet: <URL:https://arxiv.org/pdf/1702.07825 pdf pp. 1-17, (17… [cited by applicant]
Oord, A. V. D., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A . . . & Kavukcuoglu, K. (2016). Wavenet: A generative model for raw audio. arXiv preprint arXiv: 1609.03499. (Year: 2016). [cited by applicant]
Klorenzo-Trueba, J., Drugman, T., Latorre, J., Merritt, T., Putrycz, B., Barra-Chicote, R . . . & Aggarwal, V. (2018). Towards achieving robust universal neural vocoding. arXivpreprint arXiv: 1811.06292) (Year: 2018). [cited by applicant]
Ling, Z. H., Ai, Y., Gu, Y., & Dai, L. R. (2018). Waveform modeling and generation using hierarchical recurrent neural networks for speech bandwidth extension. IEEE/ACM Transactions on Audio, Speech, and Language Proces… [cited by applicant]
Liu, L. J., Ling, Z. H., Jiang, Y., Zhou, M., & Dai, L. R. (2018, September). WaveNet Vocoder with Limited Training Data for Voice Conversion. In Interspeech (pp. 1983-1987). (Year: 2018). [cited by applicant]
Mehri, S., Kumar, K., Gulrajani, I., Kumar, R., Jain, S., Sotelo, J . . . & Bengio, Y. (2016). SampleRNN: An unconditional end-to-end neural audio generation model. arXiv preprint arXiv:1612.07837. (Year: 2016). [cited by applicant]
Okamoto, T., Tachibana, K., Toda, T., Shiga, Y., & Kawai, H. (Apr. 2018). An investigation of subband WaveNet vocoder covering entire audible frequency range with limited acoustic features. In 2018 IEEE International Co… [cited by applicant]
Kalchbrenner, N., Elsen, E., Simonyan, K., Noury, S., Casagrande, N., Lockhart, E . . . & Kavukcuoglu, K. (Jul. 2018). Efficient neural audio synthesis. In International Conference on Machine Learning (pp. 2410-2419). P… [cited by applicant]
Valin, J. M., & Skoglund, J. (May 2019). LPCNet: Improving neural speech synthesis through linear prediction. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 58… [cited by applicant]
Extended European Search Report dated Mar. 30, 2022 by the European Patent Office in European Application No. 20866702.2. [cited by applicant]
Chengzhu Yu et al., “DurIAN: Duration Informed Attention Network For Multimodal Synthesis”, Sep. 5, 2019, pp. 1-11 (11 pages total), Retrieved from the Internet: URL: https://arxiv.org/pdf/1909.01700.pdf. [cited by applicant]