Adaptive enhancement of coded speech
Adaptive code enhancement includes receiving feature data of a speech waveform having a plurality of frames, inferring a bitrate based on a quantity of bits of the plurality of frames received from the decoder codec, generating one or more latent feature vectors in accordance with the bitrate and based on the feature data, receiving one or more pitch lags and a speech signal, calculating filter coefficients on a per-frame basis based on the one or more latent feature vectors and the one or more pitch lags, and modifying the speech signal using the one or more filter coefficients to generate a modified speech signal.
1 . A system, comprising:
at least one processor; and
a memory, storing program instructions that when executed by the at least one processor, cause the at least one processor to implement a deep neural network (DNN) machine learning model trained to enhance speech, the machine learning model configured to:
receive audio data decoded according to a speech codec,
determine a bitrate for the audio data based on a quantity of bits for a plurality of frames of the audio data,
extract feature data from the audio data, and
encode the feature data to produce one or more latent feature vectors according to the bitrate for the audio data,
filter the audio data according to a plurality of filter coefficients that are dynamically adjusted on a per-frame basis using the one or more latent feature vectors, and
provide the filtered audio data as speech-enhanced audio data.
2 . The system of claim 1 , wherein respective latent feature vectors of the latent features vectors utilize information within a single frame of respective feature data.
3 . The system of claim 1 , wherein the feature data comprises one or more clean speech features received by a decoder codec and one or more noisy speech features calculated from noisy decoded speech.
4 . The system of claim 1 , wherein the audio data is filtered using a plurality of filters, wherein the plurality of filters comprises:
at least one comb filter configured to:
receive the one or more latent feature vectors, one or more pitch lags, and successive versions of the audio data for filtering,
calculate respective filter coefficients for the plurality of comb filters, and
apply the respective coefficients for filter the successive versions of the audio data; and
at least one convolution filter configured to:
receive the one or more latent feature vectors and a cumulative audio data output from the plurality of comb filters, and
calculate a convolution coefficient based on the one or more latent feature vectors, and
apply the convolution coefficient to the cumulative audio data output from the plurality of comb filters to generate the speech-enhanced audio data.
5 . The system of claim 4 , wherein the convolutional filter is configured to split respective coefficients to limit amplification of the convolutional filter.
6 . The system of claim 1 , wherein the bitrate is further determined based on an exponential moving average within an update rate applied for encoding the feature data.
7 . A method, comprising:
receiving audio data decoded according to a speech codec;
processing the audio data through a deep neural network (DNN) machine learning model trained to enhance speech, wherein the DNN machine learning model:
extracts feature data from the audio data,
determines a bitrate for the audio data based on a quantity of bits received for a plurality of frames of the audio data,
encodes the feature data as one or more latent feature vectors according to the bitrate for the audio data,
filters, using at least one comb filter and at least one convolution filter, the audio data according to a plurality of filter coefficients that are dynamically adjusted on a per-frame basis using the one or more latent feature vectors, and
generates speech-enhanced audio data based, at least in part, on the filtering of the audio data.
8 . The method of claim 7 , wherein the latent feature vectors are generated at a rate that corresponds to a sub-frame rate for encoding the feature data.
9 . The method of claim 7 , wherein respective latent feature vectors of the latent features vectors utilize information within a single frame of respective features from the speech codec.
10 . The method of claim 7 , wherein the feature data comprises one or more clean speech features and one or more noisy speech features calculated from noisy decoded speech.
11 . The method of claim 7 , wherein the DNN machine learning model further normalizes respective coefficients to limit a gain for limiting amplification during filtering at the at least one convolutional filter.
12 . The method of claim 7 , wherein the bitrate is further determined based on filtering raw payload sizes for a plurality of frames of the audio data.
13 . One or more non-transitory, computer-readable storage media, storing program instructions that when executed on or across one or more computing devices of cause the one or more computing devices to implement:
receiving audio data decoded according to a speech codec;
causing the audio data to be processed through a deep neural network (DNN) machine learning model trained to enhance speech, wherein the DNN machine learning model:
extracts feature data from the audio data,
determines a bitrate for the audio data based on a quantity of bits received for a plurality of frames of the audio data,
encodes the feature data as one or more latent feature vectors according to the bitrate for the audio data,
filters the audio data according a plurality of filter coefficients that are dynamically adjusted on a per-frame basis using the one or more latent feature vectors, and
generates speech-enhanced audio data as a result of filtering the audio data.
14 . The one or more non-transitory, computer-readable storage media of claim 13 , wherein the latent feature vectors are generated at a rate that corresponds to a sub-frame rate for encoding the feature data.
15 . The one or more non-transitory, computer-readable storage media of claim 13 , wherein respective latent feature vectors of the latent features vectors utilize information within a single frame of respective features from the decoder codec.
16 . The one or more non-transitory, computer-readable storage media of claim 13 , wherein the feature data comprises one or more clean speech features and one or more noisy speech features calculated from noisy decoded speech.
17 . The one or more non-transitory, computer-readable storage media of claim 13 , wherein the bitrate is further determined based on an exponential moving average within an update rate applied for encoding the feature data.
18 . The one or more non-transitory, computer-readable storage media of claim 13 , wherein the DNN machine learning model further splits respective coefficients to limit amplification for filtering at a convolutional filter.