IP Library › Granted Patent US 12,198,710
Granted Patent B2
US 12,198,710 · App. 18/400,992 · Granted Jan 14, 2025

Generating coded data representations using neural networks and vector quantizers

Inventors: Neil Zeghidour (Paris, FR); Marco Tagliasacchi (Kilchberg, CH); Dominik Roblek (Meilen, CH)
Assignee: Google LLC
G10L19/038G06N3/045G06N3/08G10L25/30G10L2019/0002
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,710
App. No.
18/400,992
Filed
Dec 29, 2023
Granted
Jan 14, 2025
Kind
B2
Examiner
KY, KEVIN
Art Unit
2671
USPC
704/500
Abstract

Methods, systems and apparatus, including computer programs encoded on computer storage media. According to one aspect, there is provided a method comprising: receiving a new input; processing the new input using an encoder neural network to generate a feature vector representing the new input; and generating a coded representation of the feature vector using a sequence of vector quantizers that are each associated with a respective codebook of code vectors, wherein the coded representation of the feature vector identifies a plurality of code vectors, including a respective code vector from the codebook of each vector quantizer, that define a quantized representation of the feature vector.

Claims (69)

1. A method performed by one or more computers, the method comprising:

receiving a new input;

processing the new input using an encoder neural network to generate a feature vector representing the new input; and

generating a coded representation of the feature vector using a sequence of vector quantizers that are each associated with a respective codebook of code vectors,

wherein the coded representation of the feature vector identifies a plurality of code vectors, including a respective code vector from the codebook of each vector quantizer, that define a quantized representation of the feature vector, and

wherein generating the coded representation of the feature vector comprises, for a first vector quantizer in the sequence of vector quantizers:

receiving the feature vector;

identifying, based on the feature vector, a respective code vector from the codebook of the vector quantizer to represent the feature vector; and

determining a current residual vector based on an error between: (i) the feature vector, and (ii) the code vector that represents the feature vector;

wherein the coded representation of the feature vector identifies the code vector that represents the feature vector.

2. The method of claim 1 , wherein generating the coded representation of the feature vector further comprises, for each vector quantizer after the first vector quantizer in the sequence of vector quantizers:

receiving a current residual vector generated by a preceding vector quantizer in the sequence of vector quantizers;

identifying, based on the current residual vector, a respective code vector from the codebook of the vector quantizer to represent the current residual vector; and

if the vector quantizer is not a last vector quantizer in the sequence of vector quantizers:

updating the current residual vector based on an error between: (i) the current residual vector, and (ii) the code vector that represents the current residual vector;

wherein the coded representation of the feature vector identifies the code vector that represents the current residual vector.

3. The method of claim 1 , wherein the quantized representation of the feature vector is defined by a sum of the plurality of code vectors identified by the coded representation of the feature vector.

4. The method of claim 1 , wherein the codebooks of the vector quantizers all include an equal number of code vectors.

5. The method of claim 1 , wherein the encoder neural network and the codebooks of the vector quantizers have been jointly trained along with a decoder neural network, and wherein the decoder neural network is configured to:

receive a quantized representation of a feature vector representing an input; and

process the quantized representation of the feature vector representing the input to generate an output for the input.

6. The method of claim 5 , wherein the training comprised:

obtaining a plurality of training examples that each include: (i) a respective training input, and (ii) a corresponding target output;

processing the respective training input of each training example using the encoder neural network, sequence of vector quantizers, and decoder neural network to generate a respective training output that is an estimate of the corresponding target output;

determining gradients of an objective function that depends on the respective training and target outputs of each training example; and

using the gradients of the objective function to update one or more of: (i) a set of encoder neural network parameters, (ii) a set of decoder neural network parameters, or (iii) the codebooks of the vector quantizers.

7. The method of claim 6 , wherein for each training example, the corresponding target output is the same as the respective training input.

8. The method of claim 6 , wherein the objective function comprises a reconstruction loss that, for each training example, measures an error between: (i) the respective training output, and (ii) the corresponding target output.

9. The method of claim 8 , wherein the error is a mean squared error.

10. The method of claim 6 , wherein the codebooks of the vector quantizers were randomly initialized during the training.

11. The method of claim 6 , wherein the codebooks of the vector quantizers were initialized using a k-means algorithm during the training.

12. The method of claim 6 , wherein the codebooks of the vector quantizers were repeatedly updated during the training using exponential moving averages.

13. The method of claim 5 , further comprising:

processing the quantized representation of the feature vector representing the new input using the decoder neural network to generate an output for the new input.

14. A system comprising one or more computers and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

receiving a new input;

processing the new input using an encoder neural network to generate a feature vector representing the new input; and

generating a coded representation of the feature vector using a sequence of vector quantizers that are each associated with a respective codebook of code vectors,

wherein the coded representation of the feature vector identifies a plurality of code vectors, including a respective code vector from the codebook of each vector quantizer, that define a quantized representation of the feature vector, and

wherein generating the coded representation of the feature vector comprises, for a first vector quantizer in the sequence of vector quantizers:

receiving the feature vector;

identifying, based on the feature vector, a respective code vector from the codebook of the vector quantizer to represent the feature vector; and

determining a current residual vector based on an error between: (i) the feature vector, and (ii) the code vector that represents the feature vector;

wherein the coded representation of the feature vector identifies the code vector that represents the feature vector.

15. The system of claim 14 , wherein generating the coded representation of the feature vector further comprises, for each vector quantizer after the first vector quantizer in the sequence of vector quantizers:

receiving a current residual vector generated by a preceding vector quantizer in the sequence of vector quantizers;

identifying, based on the current residual vector, a respective code vector from the codebook of the vector quantizer to represent the current residual vector; and

if the vector quantizer is not a last vector quantizer in the sequence of vector quantizers:

updating the current residual vector based on an error between: (i) the current residual vector, and (ii) the code vector that represents the current residual vector;

wherein the coded representation of the feature vector identifies the code vector that represents the current residual vector.

16. The system of claim 14 , wherein the quantized representation of the feature vector is defined by a sum of the plurality of code vectors identified by the coded representation of the feature vector.

17. The system of claim 14 , wherein the codebooks of the vector quantizers all include an equal number of code vectors.

18. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

receiving a new input;

processing the new input using an encoder neural network to generate a feature vector representing the new input; and

generating a coded representation of the feature vector using a sequence of vector quantizers that are each associated with a respective codebook of code vectors,

wherein the coded representation of the feature vector identifies a plurality of code vectors, including a respective code vector from the codebook of each vector quantizer, that define a quantized representation of the feature vector, and

wherein generating the coded representation of the feature vector comprises, for a first vector quantizer in the sequence of vector quantizers:

receiving the feature vector;

identifying, based on the feature vector, a respective code vector from the codebook of the vector quantizer to represent the feature vector; and

determining a current residual vector based on an error between: (i) the feature vector, and (ii) the code vector that represents the feature vector;

wherein the coded representation of the feature vector identifies the code vector that represents the feature vector.

19. The one or more non-transitory computer storage media of claim 18 , wherein generating the coded representation of the feature vector further comprises, for each vector quantizer after the first vector quantizer in the sequence of vector quantizers:

receiving a current residual vector generated by a preceding vector quantizer in the sequence of vector quantizers;

identifying, based on the current residual vector, a respective code vector from the codebook of the vector quantizer to represent the current residual vector; and

if the vector quantizer is not a last vector quantizer in the sequence of vector quantizers:

updating the current residual vector based on an error between: (i) the current residual vector, and (ii) the code vector that represents the current residual vector;

wherein the coded representation of the feature vector identifies the code vector that represents the current residual vector.

20. The one or more non-transitory computer storage media of claim 18 , wherein the quantized representation of the feature vector is defined by a sum of the plurality of code vectors identified by the coded representation of the feature vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2024
From: ZEGHIDOUR, NEIL; TAGLIASACCHI, MARCO; ROBLEK, DOMINIK
To: GOOGLE LLC
Reel/Frame 066079/0950 →
Continuity (4)
Continuation 18106094 · Feb 6, 2023
Continuation 17856856 · Jul 1, 2022
Provisional Application 63218139 · Jul 2, 2021
Related Publication 20240185870A1 · Jun 6, 2024
References Cited (85)
US 5819215A · Dobson · 1998 [cited by examiner]
US 5966688A · Nandkumar et al. · 1999 [cited by applicant]
US 6952671B1 · Kolesnik · 2005 [cited by examiner]
US 11483646B1 · Pan · 2022 [cited by examiner]
US 11600282B2 · Zeghidour · 2023 [cited by examiner]
US 20030204394A1 · Garudadri · 2003 [cited by examiner]
US 20100023336A1 · Shmunk · 2010 [cited by examiner]
US 20150332690A1 · Kim · 2015 [cited by examiner]
US 20170251212A1 · Swaminathan · 2017 [cited by examiner]
US 20180114522A1 · Hall · 2018 [cited by examiner]
US 20180249160A1 · Zhao · 2018 [cited by examiner]
US 20190121883A1 · Swaminathan · 2019 [cited by examiner]
US 20190121884A1 · Swaminathan · 2019 [cited by examiner]
US 20200175961A1 · Thomson · 2020 [cited by examiner]
US 20200234725A1 · Garbacea et al. · 2020 [cited by applicant]
US 20220335963A1 · Jang · 2022 [cited by examiner]
US 20230186927A1 · Zeghidour · 2023 [cited by examiner]
US 20240185870A1 · Zeghidour · 2024 [cited by examiner]
Baevski et al., “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in Neural Information Processing Systems 33, 2020, 12 pages. [cited by applicant]
Barnes et al., “Advances in Residual Vector Quantization: A Review,” IEEE Transactions on Image Processing, Feb. 1996, 5(2):226-262. [cited by applicant]
Bessette et al., “The adaptive multirate wideband speech codec (AMR-WB),” IEEE Transactions on Speech and Audio Processing, Nov. 2002, 10(8):620-636. [cited by applicant]
Biswas et al., “Audio Codec Enhancement with Generative Adversarial Networks,” ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 2020, pp. 356-360. [cited by applicant]
Blau et al., “The Perception-Distortion Tradeoff,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2018, pp. 6228-6237. [cited by applicant]
Casebeer et al., “Enhancing into the Codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders,” ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Jun. 2021, … [cited by applicant]
Chinen et al., “ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric,” 2020 Twelfth International Conference on Quality of Multimedia Experience (QoMEX), May 2020, 6 pages. [cited by applicant]
Clevert et al., “Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs),” arXiv, Dec. 3, 2015, 14 pages. [cited by applicant]
Dhariwal et al., “Jukebox: A Generative Model for Music,” arXiv, Apr. 30, 2020, 20 pages. [cited by applicant]
Dieleman et al., “The challenge of realistic music generation: modelling raw audio at scale,” Advances in Neural Information Processing Systems 31, 2018, 11 pages. [cited by applicant]
Dietz et al., “Overview of the EVS codec architecture,” 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 2015, pp. 5698-5702. [cited by applicant]
Donahue et al., “Exploring Speech Enhancement with Generative Adversarial Networks for Robust Speech Recognition,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 2018, pp. 5… [cited by applicant]
Engel et al., “DDSP: Differentiable Digital Signal Processing,” arXiv, Jan. 14, 2020, 19 pages. [cited by applicant]
Feng et al., “Speech feature denoising and dereverberation via deep autoencoders for noisy reverberant speech recognition,” 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 201… [cited by applicant]
Fonseca et al., “Freesound datasets: A platform for the creation of open audio datasets,” Proceedings of the 18th ISMIR Conference, Oct. 2017, pp. 486-493. [cited by applicant]
Garbacea et al., “Low Bit-rate Speech Coding with VQ-VAE and a WaveNet Decoder,” ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 2019, pp. 735-739. [cited by applicant]
Germain et al., “Speech Denoising with Deep Feature Losses,” arXiv, Sep. 14, 2018, 6 pages. [cited by applicant]
Gray, “Vector quantization,” IEEE Assp Magazine, Apr. 1984, 1(2):4-29. [cited by applicant]
Gritsenko et al., “A Spectral Energy Distance for Parallel Speech Synthesis,” Advances in Neural Information Processing Systems 33, 2020, 11 pages. [cited by applicant]
Hines et al., “ViSQOL: The Virtual Speech Quality Objective Listener,” IWAENC 2012; International Workshop on Acoustic Signal Enhancement, Sep. 2012, 4 pages. [cited by applicant]
Holmberg et al., “Web real-time communication use cases and requirements,” IETF RFC 7478, Mar. 2015, 29 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2022/036097, mailed on Dec. 12, 2022, 18 pages. [cited by applicant]
International Telecommunication Union (ITU), “Recommendation ITU-R BS.1534-1: Method for the subjective assessment of intermediate quality level of coding systems,” Mar. 31, 2003, 18 pages. [cited by applicant]
Invitation to Pay Additional Fees in International Appln. No. PCT/US2022/036097, mailed on Oct. 19, 2022, 11 pages. [cited by applicant]
Ishii et al., “Reverberant Speech Recognition Based on Denoising Autoencoder,” Interspeech, Aug. 2013, pp. 3512-3516. [cited by applicant]
Jin et al., “FFTNet: A Real-Time Speaker-Dependent Neural Vocoder,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 2018, pp. 2251-2255. [cited by applicant]
Juang et al., “Multiple stage vector quantization for speech coding,” ICASSP '82. IEEE International Conference on Acoustics, Speech, and Signal Processing, May 1982, 7:597-600. [cited by applicant]
Kalchbrenner et al., “Efficient Neural Audio Synthesis,” Proceedings of the 35th International Conference on Machine Learning, 2018, 10 pages. [cited by applicant]
Kankanahalli et al., “End-To-End Optimized Speech Coding with Deep Neural Networks,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 2018, pp. 2521-2525. [cited by applicant]
Kleijn et al., “Generative Speech Coding with Predictive Variance Regularization,” arXiv, Feb. 18, 2021, 5 pages. [cited by applicant]
Kleijn et al., “Wavenet Based Low Rate Speech Coding,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 2018, pp. 676-680. [cited by applicant]
Kong et al., “HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,” Advances in Neural Information Processing Systems 33, 2020, 12 pages. [cited by applicant]
Kumar et al., “MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis,” Advances in Neural Information Processing Systems 32, 2019, 12 pages. [cited by applicant]
Law et al., “Evaluation of algorithms using games: The case of music tagging,” 10th International Society for Music Information Retrieval Conference (ISMIR), 2009, pp. 387-392. [cited by applicant]
Li et al., “Real-Time Speech Frequency Bandwidth Extension,” ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Jun. 2021, pp. 691-695. [cited by applicant]
Lim et al., “Time-Frequency Networks for Audio Super-Resolution,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 2018, pp. 646-650. [cited by applicant]
Linde et al., “An Algorithm for Vector Quantizer Design,” IEEE Transactions on Communications, Jan. 1980, 28(1):84-95. [cited by applicant]
Lloyd, “Least squares quantization in PCM,” IEEE Transactions on Information Theory, Mar. 1982, 28(2):129-137. [cited by applicant]
MacQueen, “Some methods for classification and analysis of multivariate observations,” Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, 1967, 1:281-297. [cited by applicant]
Makhoul et al., “Vector quantization in speech coding,” Proceedings of the IEEE, Nov. 1985, 73(11):1551-1588. [cited by applicant]
Mehri et al., “SampleRNN: An Unconditional End-to-End Neural Audio Generation Model,” Feb. 11, 2017, 11 pages. [cited by applicant]
Mentzer et al., “High-fidelity generative image compression,” Advances in Neural Information Processing Systems 33, 2020, 12 pages. [cited by applicant]
Morishima et al., “Speech coding based on a multi-layer neural network,” IEEE International Conference on Communications, Including Supercomm Technical Sessions, Apr. 1990, 2:429-433. [cited by applicant]
Pascual et al., “Segan: Speech enhancement generative adversarial network,” arXiv, Jun. 9, 2017, 5 pages. [cited by applicant]
Perez et al., “FiLM: Visual Reasoning with a General Conditioning Layer,” Proceedings of the AAAI Conference on Artificial Intelligence, Apr. 29, 2018, 32(1):3942-3951. [cited by applicant]
Polyak et al., “Speech Resynthesis from Discrete Disentangled Self-Supervised Representations,” arXiv, Jul. 27, 2021, 5 pages. [cited by applicant]
Razavi et al., “Generating Diverse High-Fidelity Images with VQ-VAE-2,” arXiv, Jun. 2, 2019, 15 pages. [cited by applicant]
Rethage et al., “A Wavenet for Speech Denoising,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 2018, pp. 5069-5073. [cited by applicant]
Schroeder et al., “Code-excited linear prediction (CELP): High-quality speech at very low bit rates,” ICASSP '85. IEEE International Conference on Acoustics, Speech, and Signal Processing, Apr. 1985, 10:937-940. [cited by applicant]
Sristava et al., “Dropout: a simple way to prevent neural networks from overfitting,” The Journal of Machine Learning Research, Jun. 2014, 15(1):1929-1958. [cited by applicant]
Stimberg et al., “WaveNetEQ—Packet Loss Concealment with WaveRNN,” 2020 54th Asilomar Conference on Signals, Systems, and Computers, Nov. 2020, pp. 672-676. [cited by applicant]
Tagliasacchi et al., “SEANet: A Multi-modal Speech Enhancement Network,” arXiv, Oct. 1, 2020, 5 pages. [cited by applicant]
Valin et al., “Definition of the Opus Audio Codec,” IETF RFC 6716, Sep. 2012, 326 pages. [cited by applicant]
Valin et al., “LPCNET: Improving Neural Speech Synthesis through Linear Prediction,” ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 2019, pp. 5891-5895. [cited by applicant]
Van den Oord et al., “Neural Discrete Representation Learning,” arXiv, Nov. 2, 2017, 10 pages. [cited by applicant]
Van den Oord et al., “Parallel WaveNet: Fast High-Fidelity Speech Synthesis,” Proceedings of the 35th International Conference on Machine Learning, Jul. 2018, 9 pages. [cited by applicant]
Van den Oord et al., “WaveNet: A Generative Model for Raw Audio,” Sep. 19, 2016, 15 pages. [cited by applicant]
Vasuki et al., “A review of vector quantization techniques,” IEEE Potentials, Jul./Aug. 2006, 25(4):39-47. [cited by applicant]
W3C.org, “WebRTC 1.0: Real-Time Communication Between Browsers,” Jan. 26, 2021, retrieved on Jul. 21, 2022, retrieved from URL <https://www.w3.org/TR/2021/REC-webrtc-20210126/>, 196 pages. [cited by applicant]
Williamson et al., “Time-Frequency Masking in the Complex Domain for Speech Dereverberation and Denoising,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, Apr. 20, 2017, 25(7):1492-1501. [cited by applicant]
Zeghidour et al., “SoundStream: An End-to-End Neural Audio Codec,” arXiv, Jul. 7, 2021, 12 pages. [cited by applicant]
Zen et al., “LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,” arXiv, Apr. 5, 2019, 7 pages. [cited by applicant]
Zhen et al., “Cascaded Cross-Module Residual Learning towards Lightweight End-to-End Speech Coding,” arXiv, Sep. 13, 2019, 5 pages. [cited by applicant]
Zhen et al., “Psychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio Coding,” IEEE Signal Processing Letters, Nov. 20, 2020, 27:2159-2163. [cited by applicant]
Chan et al., “Enhanced MultiStage Vector Quantization by Joint Codebook Design,” IEEE Transactions on Communications, Nov. 1992, 40(11):1693-1697. [cited by applicant]
Extended European Search Report in Appln. No. 24190155, mailed on Oct. 1, 2024, 12 pages. [cited by applicant]
Office Action in European Appln. 22747888.0, mailed on Sep. 25, 2024, 9 pages. [cited by applicant]