IP Library › Granted Patent US 12,744,047
Granted Patent B1
US 12,744,047 · App. 18/345,305 · Granted Sep 22, 2026

Adaptive enhancement of coded speech

Inventors: Jan Buethe (Munich, DE); Jean-Marc Valin (Montreal, CA); Ahmed Mustafa (Aachen, DE)
Assignee: Amazon Technologies, Inc.
G10L19/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,047
App. No.
18/345,305
Granted
Sep 22, 2026
Kind
B1
Abstract

Adaptive code enhancement includes receiving feature data of a speech waveform having a plurality of frames, inferring a bitrate based on a quantity of bits of the plurality of frames received from the decoder codec, generating one or more latent feature vectors in accordance with the bitrate and based on the feature data, receiving one or more pitch lags and a speech signal, calculating filter coefficients on a per-frame basis based on the one or more latent feature vectors and the one or more pitch lags, and modifying the speech signal using the one or more filter coefficients to generate a modified speech signal.

Claims (48)

1 . A system, comprising:

at least one processor; and

a memory, storing program instructions that when executed by the at least one processor, cause the at least one processor to implement a deep neural network (DNN) machine learning model trained to enhance speech, the machine learning model configured to:

receive audio data decoded according to a speech codec,

determine a bitrate for the audio data based on a quantity of bits for a plurality of frames of the audio data,

extract feature data from the audio data, and

encode the feature data to produce one or more latent feature vectors according to the bitrate for the audio data,

filter the audio data according to a plurality of filter coefficients that are dynamically adjusted on a per-frame basis using the one or more latent feature vectors, and

provide the filtered audio data as speech-enhanced audio data.

2 . The system of claim 1 , wherein respective latent feature vectors of the latent features vectors utilize information within a single frame of respective feature data.

3 . The system of claim 1 , wherein the feature data comprises one or more clean speech features received by a decoder codec and one or more noisy speech features calculated from noisy decoded speech.

4 . The system of claim 1 , wherein the audio data is filtered using a plurality of filters, wherein the plurality of filters comprises:

at least one comb filter configured to:

receive the one or more latent feature vectors, one or more pitch lags, and successive versions of the audio data for filtering,

calculate respective filter coefficients for the plurality of comb filters, and

apply the respective coefficients for filter the successive versions of the audio data; and

at least one convolution filter configured to:

receive the one or more latent feature vectors and a cumulative audio data output from the plurality of comb filters, and

calculate a convolution coefficient based on the one or more latent feature vectors, and

apply the convolution coefficient to the cumulative audio data output from the plurality of comb filters to generate the speech-enhanced audio data.

5 . The system of claim 4 , wherein the convolutional filter is configured to split respective coefficients to limit amplification of the convolutional filter.

6 . The system of claim 1 , wherein the bitrate is further determined based on an exponential moving average within an update rate applied for encoding the feature data.

7 . A method, comprising:

receiving audio data decoded according to a speech codec;

processing the audio data through a deep neural network (DNN) machine learning model trained to enhance speech, wherein the DNN machine learning model:

extracts feature data from the audio data,

determines a bitrate for the audio data based on a quantity of bits received for a plurality of frames of the audio data,

encodes the feature data as one or more latent feature vectors according to the bitrate for the audio data,

filters, using at least one comb filter and at least one convolution filter, the audio data according to a plurality of filter coefficients that are dynamically adjusted on a per-frame basis using the one or more latent feature vectors, and

generates speech-enhanced audio data based, at least in part, on the filtering of the audio data.

8 . The method of claim 7 , wherein the latent feature vectors are generated at a rate that corresponds to a sub-frame rate for encoding the feature data.

9 . The method of claim 7 , wherein respective latent feature vectors of the latent features vectors utilize information within a single frame of respective features from the speech codec.

10 . The method of claim 7 , wherein the feature data comprises one or more clean speech features and one or more noisy speech features calculated from noisy decoded speech.

11 . The method of claim 7 , wherein the DNN machine learning model further normalizes respective coefficients to limit a gain for limiting amplification during filtering at the at least one convolutional filter.

12 . The method of claim 7 , wherein the bitrate is further determined based on filtering raw payload sizes for a plurality of frames of the audio data.

13 . One or more non-transitory, computer-readable storage media, storing program instructions that when executed on or across one or more computing devices of cause the one or more computing devices to implement:

receiving audio data decoded according to a speech codec;

causing the audio data to be processed through a deep neural network (DNN) machine learning model trained to enhance speech, wherein the DNN machine learning model:

extracts feature data from the audio data,

determines a bitrate for the audio data based on a quantity of bits received for a plurality of frames of the audio data,

encodes the feature data as one or more latent feature vectors according to the bitrate for the audio data,

filters the audio data according a plurality of filter coefficients that are dynamically adjusted on a per-frame basis using the one or more latent feature vectors, and

generates speech-enhanced audio data as a result of filtering the audio data.

14 . The one or more non-transitory, computer-readable storage media of claim 13 , wherein the latent feature vectors are generated at a rate that corresponds to a sub-frame rate for encoding the feature data.

15 . The one or more non-transitory, computer-readable storage media of claim 13 , wherein respective latent feature vectors of the latent features vectors utilize information within a single frame of respective features from the decoder codec.

16 . The one or more non-transitory, computer-readable storage media of claim 13 , wherein the feature data comprises one or more clean speech features and one or more noisy speech features calculated from noisy decoded speech.

17 . The one or more non-transitory, computer-readable storage media of claim 13 , wherein the bitrate is further determined based on an exponential moving average within an update rate applied for encoding the feature data.

18 . The one or more non-transitory, computer-readable storage media of claim 13 , wherein the DNN machine learning model further splits respective coefficients to limit amplification for filtering at a convolutional filter.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2024
From: BUETHE, JAN; VALIN, JEAN-MARC; MUSTAFA, AHMED
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 066500/0129 →
References Cited (29)
US 9741351B2 · Vinton · 2017 [cited by examiner]
US 11521637B1 · Valin et al. · 2022 [cited by applicant]
US 20120140938A1 · Yoo · 2012 [cited by examiner]
US 20240274144A1 · Shi · 2024 [cited by examiner]
US 20250014584A1 · Pia · 2025 [cited by examiner]
J.-H. Chen and A. Gersho, “Adaptive Postfiltering for Quality Enhancement of Coded Speech,” IEEE Transactions on Speech and Audio Processing, vol. 3, No. 1, pp. 59-71, 1995. [cited by applicant]
Z. Zhao, H. Liu, and T. Fingscheidt, “Convolutional Neural Networks to Enhance Coded Speech,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 27, No. 4, pp. 663-678, 2019. [cited by applicant]
J. Skoglund and J.-M. Valin, “Improving Opus Low Bit Rate Quality with Neural Speech Synthesis,” in Proc. INTERSPEECH, 2019, pp. 2847-2851. [cited by applicant]
K. Gupta, S. Korse, B. Edler, and G. Fuchs, “A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT Domain,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, a… [cited by applicant]
S. Korse, K. Gupta, and G. Fuchs, “Enhancement of Coded Speech Using a Mask-Based Post-Filter,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, arXiv:2010.05571v1, pp. 1-5. [cited by applicant]
S. Korse, N. Pia, K. Gupta, and G. Fuchs, “PostGAN: A GAN-Based Post-Processor to Enhance the Quality of Coded Speech,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, arXiv:… [cited by applicant]
P. Ochieng, “Deep Neural Network Techniques for Monaural Speech Enhancement: State of the Art Analysis,” arXiv:2212.00369, 2022, pp. 1-46. [cited by applicant]
J -M. Valin, K. Vos, and T. B. Terriberry, “Definition of the Opus Audio Codec,” RFC 6716, 2012, 1 page. [cited by applicant]
J.-M. Valin and J. Skoglund, “A Real-Time Wideband Neural Vocoder at 1.6kb/s Using LPCNet,” in Proc. INTERSPEECH, 2019, pp. 3406-3410. [cited by applicant]
W. B. Kleijn, A. Storus, M. Chinen, T. Denton, F. S. C. Lim, A. Luebs, J. Skoglund, and H. Yeh, “Generative Speech Coding with Predictive Variance Regularization,” in IEEE Proc. International Conference on Acoustics, Sp… [cited by applicant]
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “SoundStream: An End-to-End Neural Audio Codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, pp. 495-507, 2022 (ArXiv pre… [cited by applicant]
N. Pia, K. Gupta, S. Korse, M. Multrus, and G. Fuchs, “NESC: Robust Neural End-2-End Speech Coding with GANs,” in Proc. INTERSPEECH, 2022, arXiv:2207.03282v1, pp. 1-5. [cited by applicant]
T. Jenrungrot, M. Chinen, W. Kleijn, J. Skoglund, Z. Borsos, N. Zeghidour, and M. Tagliasacchi, “Lmcodec: A low bitrate speech codec with causal transformer models,” in Proc. International Conference on Acoustics, Speec… [cited by applicant]
T. Salimans and D. P. Kingma, “Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks,” 30th Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain, p… [cited by applicant]
I. Demirsahin, O. Kjartansson, A. Gutkin, and C. Rivera, “Open-source Multi-speaker Corpora of the English Accents in the British Isles,” in Proceedings of the 12th Conference on Language Resources and Evaluation (LREC … [cited by applicant]
O. Kjartansson, A. Gutkin, A. Butryna, I. Demirsahin, and C. Rivera, “Open-Source High Quality Speech Datasets for Basque, Catalan and Galician,” in Proceedings of the 1st Joint SLTU and CCURL Workshop (SLTU-CCURL 2020)… [cited by applicant]
K. Sodimana, K. Pipatsrisawat, L. Ha, M. Jansche, O. Kjartansson, P. D. Silva, and S. Sarin, “A Step-by-Step Process for Building TTS Voices Using Open Source Data and Framework for Bangla, Javanese, Khmer, Nepali, Sinh… [cited by applicant]
A. Guevara-Rukoz, I. Demirsahin, F. He, S.-H. C. Chu, S. Sarin, K. Pipatsrisawat, A. Gutkin, A. Butryna, and O. Kjartansson, “Crowdsourcing Latin American Spanish for Low-Resource Text-to-Speech,” in Proceedings of the … [cited by applicant]
Y. M. Oo, T. Wattanavekin, C. Li, P. De Silva, S. Sarin, K. Pipatsrisawat, M. Jansche, O. Kjartansson, and A. Gutkin, “Burmese Speech Corpus, Finite-State Text Normalization and Pronunciation Grammars with an Applicatio… [cited by applicant]
D. van Niekerk, C. van Heerden, M. Davel, N. Kleynhans, O. Kjartansson, M. Jansche, and L. Ha, “Rapid development of TTS corpora for four South African languages,” 2017 ISCA, in Proc. INTERSPEECH, 2017, pp. 2178-2182. [cited by applicant]
A. Gutkin, I. Demirs, ahin, O. Kjartansson, C. Rivera, and K. T'ub'o. s'un, “Developing an Open-Source Corpus of Yoruba Speech,” in Proc. INTERSPEECH, 2020 ISCA, pp. 404-408. [cited by applicant]
E. Bakhturina, V. Lavrukhin, B. Ginsburg, and Y. Zhang, “Hi-Fi Multi-Speaker English TTS Dataset,” in Proc. INTERSPEECH, 2021, pp. 2776-2780 (retrieved from arXiv:2104.01497v3, 2021 pp. 1-5). [cited by applicant]
B. Moore, “An Introduction to the Psychology of Hearing”. Brill, Sixth Edition, pp. 1-456, 2013. [cited by applicant]
K. Vos, K. Sorensen, S. Jensen, and J.-M. Valin, “Voice Coding with Opus,” 135th AES Convention, pp. 722-731, 2013. [cited by applicant]