IP Library › Granted Patent US 12,518,773
Granted Patent B2
US 12,518,773 · App. 18/278,746 · Granted Jan 6, 2026

Compressing audio waveforms using a structured latent space

Inventors: Ahmed Omran (Baar, CH); Neil Zeghidour (Paris, FR); Zalán Borsos (Zurich, CH); Félix de Chaumont Quitry (Zürich, CH); Marco Tagliasacchi (Kilchberg, CH)
Assignee: Google LLC
G10L21/0208G10L19/038
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,773
App. No.
18/278,746
Filed
Aug 24, 2023
Granted
Jan 6, 2026
Kind
B2
Art Unit
2655
USPC
704/500
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training an encoder neural network and a decoder neural network. In one aspect, a method includes obtaining a first initial audio waveform and a first noisy audio waveform, obtaining a second initial audio waveform and a second noisy audio waveform, processing the first noisy audio waveform and the second noisy audio waveform using an encoder neural network, generating a blended embedding by concatenating: (i) clean feature dimensions from an embedding of the first noisy audio waveform, and (ii) noise feature dimensions from an embedding of the second noisy audio waveform, processing the blended embedding using a decoder neural network to generate a reconstructed audio waveform, determining gradients of an objective function; and updating parameter values of the encoder neural network and the decoder neural network using the gradients.

Claims (112)

1 . A method performed by one or more computers, the method comprising:

obtaining a first initial audio waveform and a first noisy audio waveform, wherein the first noisy audio waveform is generated by applying a first set of noise parameters to the first initial audio waveform;

obtaining a second initial audio waveform and a second noisy audio waveform, wherein the second noisy audio waveform is generated by applying a second set of noise parameters to the second initial audio waveform;

processing the first noisy audio waveform and the second noisy audio waveform using an encoder neural network, wherein:

the encoder neural network is configured to process an input audio waveform to generate embedding of the input audio waveform; and

the embedding of the input audio waveform comprises a plurality of feature dimensions, wherein the plurality of feature dimensions comprises: (i) a set of feature dimensions designated as clean feature dimensions, and (ii) a set of feature dimensions designated as noise feature dimensions;

generating a blended embedding by concatenating: (i) clean feature dimensions from the embedding of the first noisy audio waveform, and (ii) noise feature dimensions from the embedding of the second noisy audio waveform;

processing the blended embedding using a decoder neural network to generate a reconstructed audio waveform;

determining gradients of an objective function that measures an error between: (i) the reconstructed audio waveform, and (ii) an audio waveform generated by applying the second set of noise parameters to the first initial audio waveform; and

updating parameter values of the encoder neural network and the decoder neural network using the gradients.

2 . The method of claim 1 , wherein the first set of noise parameters comprises a first noise waveform.

3 . The method of claim 2 , wherein applying the first set of noise parameters to the first initial audio waveform comprises:

adding the first noise waveform to the first initial audio waveform.

4 . The method of claim 2 , wherein applying the first set of noise parameters to the first initial audio waveform comprises:

convolving the first noise waveform with the first initial audio waveform.

5 . The method of claim 1 , wherein the objective function measures the error between: (i) the reconstructed audio waveform, and (ii) the audio waveform generated by applying the second set of noise parameters to the first initial audio waveform, by a multi-scale spectral reconstruction loss.

6 . The method of claim 1 , wherein the set of feature dimensions designated as clean feature dimensions is disjoint from the set of feature dimensions designated as noise feature dimensions.

7 . The method of claim 1 , wherein the embedding of the input audio waveform comprises a plurality of feature vectors representing the input audio waveform, wherein each feature vector comprises: (i) the set of feature dimensions designated as clean feature dimensions, and (ii) the set of feature dimensions designated as noise feature dimensions.

8 . The method of claim 1 , wherein generating the blended embedding comprises vector quantizing the blended embedding.

9 . The method of claim 8 , wherein vector quantizing the blended embedding comprises:

vector quantizing the clean feature dimensions of the blended embedding using a first vector quantizer; and

vector quantizing the noise feature dimensions of the blended embedding using a second vector quantizer.

10 . The method of claim 1 , wherein the first initial audio waveform and the second initial audio waveform are speech waveforms or music waveforms.

11 . The method of claim 1 , wherein the encoder neural network and the decoder neural network have respective convolutional neural network architectures.

12 . The method of claim 1 , further comprising:

processing data derived from the reconstructed audio waveform using a discriminator neural network to generate a set of one or more discriminator scores, wherein each discriminator score characterizes an estimated likelihood that the reconstructed audio waveform was generated using the encoder neural network and the decoder neural network;

wherein the objective function further comprises an adversarial loss that depends on the discriminator scores generated by the discriminator neural network.

13 . The method of claim 12 , wherein the objective function measures an error between: (i) one or more intermediate outputs generated by the discriminator neural network by processing the reconstructed audio waveform, and (ii) one or more intermediate outputs generated by the discriminator neural network by processing the audio waveform generated by applying the second set of noise parameters to the first initial audio waveform.

14 . The method of claim 1 , further comprising:

obtaining a third initial audio waveform and a third noisy audio waveform, wherein the third noisy audio waveform is generated by applying a third set of noise parameters to the third initial audio waveform;

processing the third noisy audio waveform using the encoder neural network to generate an embedding of the third noisy audio waveform;

generating a clean embedding by setting values of the noise feature dimensions of the embedding of the third noisy audio waveform to default values;

processing the clean embedding using the decoder neural network to generate a reconstructed audio waveform;

determining gradients of an objective function that measures an error between: (i) the reconstructed audio waveform, and (ii) the third initial audio waveform; and

updating parameter values of the encoder neural network and the decoder neural network using the gradients.

15 . The method of claim 14 , wherein generating the clean embedding by setting values of the noise feature dimensions of the embedding of the third noisy audio waveform to default values comprises:

setting values of the noise feature dimensions of the embedding of the third noisy audio waveform to zero.

16 . The method of claim 1 , further comprising:

obtaining a fourth audio waveform;

processing the fourth audio waveform using the encoder neural network to generate an embedding of the fourth audio waveform;

processing the embedding of the fourth audio waveform to generate a reconstructed audio waveform;

determining gradients of an objective function that measures an error between: (i) the reconstructed audio waveform, and (ii) the fourth audio waveform; and

updating parameter values of the encoder neural network and the decoder neural network using the gradients.

17 . The method of claim 1 , wherein determining gradients of an objective function that measures an error between: (i) the reconstructed audio waveform, and (ii) an audio waveform generated by applying the second set of noise parameters to the first initial audio waveform, comprises:

backpropagating gradients of the objective function through the decoder neural network and into the encoder neural network.

18 . The method of claim 1 , wherein updating parameter values of the encoder neural network and the decoder neural network using the gradients comprises:

updating the parameter values of the encoder neural network and the decoder neural network using the gradients in accordance with a gradient descent optimization technique.

19 . A method performed by one or more computers, the method comprising:

obtaining an audio waveform;

processing the audio waveform using an encoder neural network to generate an embedding of the audio waveform;

wherein the encoder neural network has been trained by performing operations comprising:

obtaining a first initial audio waveform and a first noisy audio waveform, wherein the first noisy audio waveform is generated by applying a first set of noise parameters to the first initial audio waveform;

obtaining a second initial audio waveform and a second noisy audio waveform, wherein the second noisy audio waveform is generated by applying a second set of noise parameters to the second initial audio waveform;

processing the first noisy audio waveform and the second noisy audio waveform using the encoder neural network, wherein:

the encoder neural network is configured to process an input audio waveform to generate embedding of the input audio waveform; and

the embedding of the input audio waveform comprises a plurality of feature dimensions, wherein the plurality of feature dimensions comprises: (i) a set of feature dimensions designated as clean feature dimensions, and (ii) a set of feature dimensions designated as noise feature dimensions;

generating a blended embedding by concatenating: (i) clean feature dimensions from the embedding of the first noisy audio waveform, and (ii) noise feature dimensions from the embedding of the second noisy audio waveform;

processing the blended embedding using a decoder neural network to generate a reconstructed audio waveform;

determining gradients of an objective function that measures an error between: (i) the reconstructed audio waveform, and (ii) an audio waveform generated by applying the second set of noise parameters to the first initial audio waveform; and

updating parameter values of the encoder neural network and the decoder neural network using the gradients;

vector quantizing the embedding of the audio waveform; and

compressing the quantized embedding of the audio waveform.

20 . The method of claim 19 , further comprising, before compressing the quantized representation of the audio waveform:

removing noise feature dimensions of the embedding of the audio waveform.

21 . The method of claim 19 , further comprising, before compressing the quantized representation of the audio waveform:

scaling noise feature dimensions of the embedding of the audio waveform.

22 . The method of claim 19 , wherein compressing the quantized embedding of the audio waveform comprises:

compressing clean feature dimensions of the quantized embedding of the audio waveform at a higher bit rate than noise feature dimensions of the quantized embedding of the audio waveform.

23 . The method of claim 19 , wherein compressing the quantized embedding of the audio waveform comprises:

compressing the quantized embedding of the audio waveform using entropy encoding techniques.

24 . A method performed by one or more computers, the method comprising:

receiving a compressed quantized embedding of an audio waveform that is generated by the performing operations comprising:

processing the audio waveform using an encoder neural network to generate an embedding of the audio waveform;

wherein the encoder neural network has been trained by performing operations comprising:

obtaining a first initial audio waveform and a first noisy audio waveform, wherein the first noisy audio waveform is generated by applying a first set of noise parameters to the first initial audio waveform;

obtaining a second initial audio waveform and a second noisy audio waveform, wherein the second noisy audio waveform is generated by applying a second set of noise parameters to the second initial audio waveform;

processing the first noisy audio waveform and the second noisy audio waveform using the encoder neural network, wherein:

the encoder neural network is configured to process an input audio waveform to generate embedding of the input audio waveform; and

the embedding of the input audio waveform comprises a plurality of feature dimensions, wherein the plurality of feature dimensions comprises: (i) a set of feature dimensions designated as clean feature dimensions, and (ii) a set of feature dimensions designated as noise feature dimensions;

generating a blended embedding by concatenating: (i) clean feature dimensions from the embedding of the first noisy audio waveform, and (ii) noise feature dimensions from the embedding of the second noisy audio waveform;

processing the blended embedding using a decoder neural network to generate a reconstructed audio waveform;

determining gradients of an objective function that measures an error between: (i) the reconstructed audio waveform, and (ii) an audio waveform generated by applying the second set of noise parameters to the first initial audio waveform; and

updating parameter values of the encoder neural network and the decoder neural network using the gradients;

vector quantizing the embedding of the audio waveform; and

compressing the quantized embedding of the audio waveform; and

decompressing the compressed quantized embedding of the audio waveform.

25 . A method performed by one or more computers, the method comprising:

obtaining a compressed quantized embedding of an audio waveform;

decompressing the compressed quantized embedding of the audio waveform, wherein the quantized embedding of the audio waveform comprises: (i) a set of vector quantized feature dimensions designated as clean feature dimensions that represent an initial audio signal in the audio waveform, and (ii) a set of vector quantized feature dimensions designated as noise feature dimensions that represent a noisy audio signal in the audio waveform; and

processing the quantized embedding of the audio waveform using a decoder neural network to generate a reconstruction of the audio waveform.

26 . A system comprising:

one or more computers; and

one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

obtaining a first initial audio waveform and a first noisy audio waveform, wherein the first noisy audio waveform is generated by applying a first set of noise parameters to the first initial audio waveform;

obtaining a second initial audio waveform and a second noisy audio waveform, wherein the second noisy audio waveform is generated by applying a second set of noise parameters to the second initial audio waveform;

processing the first noisy audio waveform and the second noisy audio waveform using an encoder neural network, wherein:

the encoder neural network is configured to process an input audio waveform to generate embedding of the input audio waveform; and

the embedding of the input audio waveform comprises a plurality of feature dimensions, wherein the plurality of feature dimensions comprises: (i) a set of feature dimensions designated as clean feature dimensions, and (ii) a set of feature dimensions designated as noise feature dimensions;

generating a blended embedding by concatenating: (i) clean feature dimensions from the embedding of the first noisy audio waveform, and (ii) noise feature dimensions from the embedding of the second noisy audio waveform;

processing the blended embedding using a decoder neural network to generate a reconstructed audio waveform;

determining gradients of an objective function that measures an error between: (i) the reconstructed audio waveform, and (ii) an audio waveform generated by applying the second set of noise parameters to the first initial audio waveform; and

updating parameter values of the encoder neural network and the decoder neural network using the gradients.

27 . One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

obtaining a first initial audio waveform and a first noisy audio waveform, wherein the first noisy audio waveform is generated by applying a first set of noise parameters to the first initial audio waveform;

obtaining a second initial audio waveform and a second noisy audio waveform, wherein the second noisy audio waveform is generated by applying a second set of noise parameters to the second initial audio waveform;

processing the first noisy audio waveform and the second noisy audio waveform using an encoder neural network, wherein:

the encoder neural network is configured to process an input audio waveform to generate embedding of the input audio waveform; and

the embedding of the input audio waveform comprises a plurality of feature dimensions, wherein the plurality of feature dimensions comprises: (i) a set of feature dimensions designated as clean feature dimensions, and (ii) a set of feature dimensions designated as noise feature dimensions;

generating a blended embedding by concatenating: (i) clean feature dimensions from the embedding of the first noisy audio waveform, and (ii) noise feature dimensions from the embedding of the second noisy audio waveform;

processing the blended embedding using a decoder neural network to generate a reconstructed audio waveform;

determining gradients of an objective function that measures an error between: (i) the reconstructed audio waveform, and (ii) an audio waveform generated by applying the second set of noise parameters to the first initial audio waveform; and

updating parameter values of the encoder neural network and the decoder neural network using the gradients.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2023
From: OMRAN, AHMED; ZEGHIDOUR, NEIL; BORSOS, ZALAN; DE CHAUMONT QUITRY, FELIX; TAGLIASACCHI, MARCO
To: GOOGLE LLC
Reel/Frame 065235/0190 →
Continuity (2)
Provisional Application 63321407 · Mar 18, 2022
Related Publication 20250022477A1 · Jan 16, 2025
References Cited (44)
US 11257507B2 · Garbacea · 2022 [cited by examiner]
US 11600282B2 · Zeghidour · 2023 [cited by examiner]
US 20200349965A1 · Nesta · 2020 [cited by examiner]
US 20210166706A1 · Lim · 2021 [cited by examiner]
US 20230267940A1 · Jang · 2023 [cited by examiner]
US 20230274141A1 · Sung · 2023 [cited by examiner]
US 20230335109A1 · Zhang · 2023 [cited by examiner]
US 20240095499A1 · Lee · 2024 [cited by examiner]
US 20250022477A1 · Omran · 2025 [cited by examiner]
Trinh, Viet Anh, and Sebastian Braun. “Unsupervised speech enhancement with speech recognition embedding and disentanglement losses.” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Proces… [cited by examiner]
Bhagat et al., “DisCont: Self-supervised visual attribute disentanglement using context vectors,” CoRR, submitted on Jun. 29, 2020, arXiv:2006.05895v2, 11 pages. [cited by applicant]
Chinen et al., “ViSQOL v3: An open source production ready objective speech and audio metric,” Proceedings of International Conference on Quality of Multimedia Experience (QoMEX), May 26, 2020, pp. 1-6. [cited by applicant]
Chou et al., “Multi-target voice conversion without parallel data by adversarially learning disentangled audio representations,” CoRR, submitted on Jun. 24, 2018, arXiv:1804.02812v2, 6 pages. [cited by applicant]
Engel et al., “DDSP: Differentiable digital signal processing,” CoRR, submitted on Jan. 14, 2020, arXiv:2001.04643v1, 19 pages. [cited by applicant]
Fonseca et al., “Freesound datasets: a platform for the creation of open audio datasets,” Proceedings of the 18th International Society for Music Information Retrieval Conference, Oct. 23-27, 2017, pp. 486-493. [cited by applicant]
Garbacea et al., “Low bit-rate speech coding with VQ-VAE and a WaveNet decoder,” Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 16, 2019, pp. 735-739. [cited by applicant]
Gfeller et al., “SPICE: Self-supervised pitch estimation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, Mar. 20, 2020, 28:1118-1128. [cited by applicant]
Gong et al., “Towards learning fine-grained disentangled representations from speech,” CoRR, submitted on Aug. 8, 2018, arXiv:1808.02939v1, 3 pages. [cited by applicant]
Gritsenko et al., “A spectral energy distance for parallel speech synthesis,” Proceedings of the 34th International Conference on Neural Information Processing Systems, Dec. 2020, 33:13062-13072. [cited by applicant]
Hajihassnai et al., “ObscureNet: Learning attribute-invariant latent representation for anonymizing sensor data,” Proceedings of International Conference on Internet-of-Things Design and Implementation, May 18, 2021, pp… [cited by applicant]
Higgins et al., “Towards a definition of disentangled representations,” CoRR, submitted on Dec. 5, 2018, arXiv:1812.02230v1, 29 pages. [cited by applicant]
Hines et al., “ViSQOL: An objective speech quality model,” EURASIP Journal on Audio, Speech, and Music Processing, Dec. 2015, 18 pages. [cited by applicant]
Hu et al., “Disentangling factors of variation by mixing them,” Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 18-22, 2018, pp. 3399-3407. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2023/015392, mailed on Jul. 10, 2023, 17 pages. [cited by applicant]
Kankanahalli, “End-To-end optimized speech coding with deep neural networks,” CoRR, submitted on Dec. 9, 2019, arXiv:1710.09064v2, 5 pages. [cited by applicant]
Kavalerov et al., “Universal sound separation,” Proceedings of IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), Oct. 20-23, 2019, pp. 175-179. [cited by applicant]
Kulkarni et al., “Deep convolutional inverse graphics network,” Proceedings of the 28th International Conference on Neural Information Processing Systems, Dec. 2015, 28:2539-2547. [cited by applicant]
Lample et al., “Fader networks: Manipulating images by sliding attributes,” Proceedings of Advances in Neural Information Processing Systems, Dec. 2017, 30:5969-5978. [cited by applicant]
Locatello et al., “A sober look at the unsupervised learning of disentangled representations and their evaluation,” CoRR, submitted on Oct. 27, 2020, arXiv:2010.14766v1, 62 pages. [cited by applicant]
Mor et al., “A universal music translation network,” CoRR, submitted on May 23, 2018, arXiv:1805.07848v2, 10 pages. [cited by applicant]
Morishima et al., “Speech coding based on a multi-layer neural network,” Proceedings of IEEE International Conference on Communications, vol. 2, Apr. 16, 1990, pp. 429-433. [cited by applicant]
Nguyen et al., “NVC-Net: End-to-end adversarial voice conversion,” CoRR, submitted on Jun. 2, 2021, arXiv:2106.00992v1, 16 pages. [cited by applicant]
O'Malley et al., “A Conformer-based ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement and Speech Separation,” CoRR, submitted on Nov. 18, 2021, arXiv:2111.09935v1, 8 pages. [cited by applicant]
Park et al., “Swapping autoencoder for deep image manipulation,” Proceedings of Advances in Neural Information Processing Systems, Dec. 2020, 33:7198-7211. [cited by applicant]
Polyak et al., “Speech resynthesis from discrete disentangled self-supervised representations,” CoRR, submitted on Jul. 27, 2021, arXiv:2104.00355v3, 5 pages. [cited by applicant]
Qian et al., “AutoVC: Zero-shot voice style transfer with only autoencoder loss,” Proceedings of International Conference on Machine Learning, May 24, 2019, pp. 5210-5219. [cited by applicant]
Qian et al., “Unsupervised speech decomposition via triple information bottleneck,” Proceedings of International Conference on Machine Learning, Nov. 21, 2020, pp. 7836-7846. [cited by applicant]
Valin et al., “Definition of the Opus audio codec,” Internet Engineering Task Force, Request for Comments RFC 6716, Sep. 2012, 326 pages. [cited by applicant]
Wisdom et al., “Unsupervised sound separation using mixture invariant training,” Proceedings of Advances in Neural Information Processing Systems, Dec. 2020, 33:3846-3857. [cited by applicant]
Xie et al., “Noisy-to-Noisy Voice Conversion Framework with Denoising Model,” CoRR, submitted on Sep. 22, 2021, arXiv:2109.10608v1, 7 pages. [cited by applicant]
Yang et al., “Source-aware neural speech coding for noisy speech compression,” Proceedings of 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 13, 2021, pp. 706-710. [cited by applicant]
Zeghidour et al., “SoundStream: An end-to-end neural audio codec,” CoRR, submitted on Jul. 7, 2021, arXiv:2107.03312v1, 12 pages. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2023/015392, mailed on Oct. 3, 2024, 11 pages. [cited by applicant]
Office Action in European Appln. No. 23716724.2, mailed on Jun. 17, 2025, 6 pages. [cited by applicant]