IP Library › Granted Patent US 12,205,605
Granted Patent B2
US 12,205,605 · App. 17/670,172 · Granted Jan 21, 2025

Audio signal encoding and decoding method using a neural network model to generate a quantized latent vector, and encoder and decoder for performing the same

Inventors: Inseon Jang (Daejeon, KR); Seung Kwon Beack (Daejeon, KR); Jongmo Sung (Daejeon, KR); Tae Jin Lee (Daejeon, KR); Woo-Taek Lim (Sejong, KR); Hong-Goo Kang (Seoul, KR); Jihyun Lee (Seoul, KR); Chanwoo Lee (Seoul, KR); Hyungseob Lim (Seoul, KR)
Assignees: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE; INDUSTRY-ACADEMIC COOPERATION FOUNDATION, YONSEI UNIVERSITY
G10L19/038G10L25/30G10L2019/0001
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,605
App. No.
17/670,172
Granted
Jan 21, 2025
Kind
B2
Abstract

An audio signal encoding and decoding method using a neural network model, and an encoder and decoder for performing the same are disclosed. A method of encoding an audio signal using a neural network model, the method may include identifying an input signal, generating a quantized latent vector by inputting the input signal into a neural network model encoding the input signal, and generating a bitstream corresponding to the quantized latent vector, wherein the neural network model may include i) a feature extraction layer generating a latent vector by extracting a feature of the input signal, ii) a plurality of downsampling blocks downsampling the latent vector, and iii) a plurality of quantization blocks performing quantization of a downsampled latent vector.

Claims (25)

1. A method of encoding an audio signal using a neural network model, the method comprising:

identifying an input signal;

generating a quantized latent vector by inputting the input signal into the neural network model encoding the input signal; and

generating a bitstream corresponding to the quantized latent vector,

wherein the neural network model comprises i) a feature extraction layer generating a latent vector by extracting a feature of the input signal, ii) a plurality of downsampling blocks producing a plurality of downsampled latent vectors corresponding to downsamplings of the latent vector respectively, and iii) a plurality of quantization blocks performing quantization of the plurality of downsampled latent vectors respectively, and

wherein each of the plurality of quantization blocks comprises a conversion layer converting the respective downsampled latent vector to produce a respective converted latent vector, and a vector quantization layer performing vector quantization on the respective converted latent vector based on a codebook.

2. The method of claim 1 , wherein the plurality of downsampled latent vectors respectively correspond to downsamplings of the quantized latent vector to different respective time resolutions.

3. The method of claim 1 , wherein the vector quantization layer performs vector quantization of the converted latent vector by determining a code in the codebook in a nearest distance from the converted latent vector.

4. The method of claim 1 , wherein a downsampling block of the plurality of downsampling blocks comprises a convolution layer performing a convolution operation and a maxpool layer processing a max-pooling operation on an operation result of the convolution layer.

5. The method of claim 4 , wherein the downsampling block further comprises a residual block increasing non-linearity of an operation result of the maxpool layer, and the residual block comprises a convolution layer performing a convolution operation, a batch normalization layer performing batch normalization, and an activation layer.

6. A method of decoding an audio signal using a neural network model, the method comprising:

identifying a bitstream generated by an encoder; and

generating an output signal by inputting the bitstream into the neural network model generating an output signal from the bitstream,

wherein the neural network model comprises a plurality of inverse-quantization blocks extracting respective inverse-quantized latent vectors having different respective time resolutions from the bitstream, a plurality of upsampling blocks upsampling the inversely-quantized latent vectors to produce upsampled latent vectors, respectively, and a restoration layer generating an output signal from the upsampled latent vectors.

7. The method of claim 6 , wherein the plurality of upsampling blocks upsamples the inversely-quantized latent vectors in an ascending order of time resolutions, and a current upsampling block of the plurality of upsampling blocks upsamples a combination of i) a latent vector, among the inversely-quantized latent vectors, having a same time resolution as an upsampled latent vector produced by a previous upsampling block of the plurality of upsampling blocks and ii) the upsampled latent vector produced by the previous upsampling block.

8. The method of claim 6 , wherein the inverse-quantization block comprises a residual block increasing non-linearity of the latent vector, and a convolution layer performing a convolution operation.

9. An encoder for performing a method of encoding an audio signal using a neural network model, the encoder comprising:

a processor configured to identify an input signal, generate a quantized latent vector by inputting the input signal to the neural network model encoding the input signal, and generate a bitstream corresponding to the quantized latent vector,

wherein the neural network model comprises i) a feature extraction layer generating a latent vector by extracting a feature of the input signal, ii) a plurality of downsampling blocks producing a plurality of downsampled latent vector corresponding to downsamplings of the latent vector respectively, and iii) a plurality of quantization blocks performing quantization of the plurality of downsampled latent vectors respectively, and

wherein each of the plurality of quantization blocks comprises a conversion layer converting the respective downsampled latent vector to produce a respective converted latent vector, and a vector quantization layer performing vector quantization on the respective converted latent vector based on a codebook.

10. The encoder of claim 9 , wherein the plurality of downsampled latent vectors respectively correspond to downsamplings of the quantized latent vector to different respective time resolutions.

11. The encoder of claim 9 , wherein the vector quantization layer performs vector quantization of the converted latent vector by determining a code in the codebook in a nearest distance from the converted latent vector.

12. The encoder of claim 9 , wherein a downsampling block of the plurality of downsampling blocks comprises a convolution layer performing a convolution operation and a maxpool layer processing a max-pooling operation on an operation result of the convolution layer.

13. The encoder of claim 12 , wherein the downsampling block further comprises a residual block increasing non-linearity of an operation result of the maxpool layer, and

the residual block comprises a convolution layer performing a convolution operation, a batch normalization layer performing batch normalization, and an activation layer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2022
From: JANG, INSEON; BEACK, SEUNG KWON; SUNG, JONGMO; LEE, TAE JIN; LIM, WOO-TAEK; KANG, HONG-GOO; LEE, JIHYUN; LEE, CHANWOO; LIM, HYUNGSEOB
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE; INDUSTRY-ACADEMIC COOPERATION FOUNDATION, YONSEI UNIVERSITY
Reel/Frame 058999/0804 →
Priority Claims (1)
KR 10-2021-0049104 · Apr 15, 2021 · national
Continuity (1)
Related Publication 20220335963A1 · Oct 20, 2022
References Cited (17)
US 7630902B2 · You · 2009 [cited by examiner]
US 10068557B1 · Engel · 2018 [cited by examiner]
US 10665247B2 · Vasilache · 2020 [cited by examiner]
US 11257507B2 · Garbacea · 2022 [cited by examiner]
US 20180249160A1 · Zhao · 2018 [cited by examiner]
US 20190164052A1 · Sung et al. · 2019 [cited by applicant]
US 20200135220A1 · Lee · 2020 [cited by examiner]
US 20230267600A1 · Qu · 2023 [cited by examiner]
US 20230377584A1 · Pascual · 2023 [cited by examiner]
KR 1020200039530A · 2020 [cited by applicant]
Chen, Mingjie, and Thomas Hain. “Unsupervised acoustic unit representation learning for voice conversion using wavenet auto-encoders.” arXiv preprint arXiv:2008.06892 (2020) (Year: 2020). [cited by examiner]
Zhen, Kai, et al. “Efficient and scalable neural residual waveform coding with collaborative quantization.” ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020.… [cited by examiner]
Fostiropoulos, Iordanis. “Depthwise Discrete Representation Learning.” arXiv preprint arXiv:2004.05462 (Year: 2020). [cited by examiner]
Wang, Xin, et al. “A Vector Quantized Variational Autoencoder (VQ-VAE) Autoregressive Neural F0 Model for Statistical Parametric Speech Synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing 28, 201… [cited by applicant]
Engel, Jesse, et al. “Neural audio synthesis of musical notes with wavenet autoencoders,” International Conference on Machine Learning, PMLR, 2017. [cited by applicant]
Mingjie Chen et al., “Unsupervised Acoustic Unit Representation Learning for Voice Conversion using WaveNet Auto-encoders,” Interspeech 2020, Aug. 16, 2020, arXiv preprint arXiv:2008.06892. [cited by applicant]
Siang Thye Hang et al.. “Bi-linearly weighted fractional max pooling: An extension to conventional max pooling for deep convolutional neural network.” Multimedia Tools and Applications, May 31, 2017. [cited by applicant]