NOISE SUPPRESSION SYSTEM AND METHOD
Systems and methods are described for applying noise suppression to one or more audio signals to generate a noise-suppressed audio signal therefrom. In a single-channel implementation, an input signal is received that comprises a desired audio signal and an additive noise signal. Noise suppression is then applied to the input signal to generate a noise-suppressed signal in a manner that is controlled by at least a parameter that specifies a degree of balance between distortion of the desired audio signal and unnaturalness of a residual noise signal included in the noise-suppressed signal. In an alternative single-channel implementation, a plurality of sub-band signals obtained by applying a frequency conversion process to a time domain representation of an input signal is received. Noise suppression is then applied to each of the sub-band signals by passing each of the sub-band signals through a time direction filter. Multi-channel noise suppression variants are also described.
1 . A method, comprising:
receiving an input audio signal that comprises a desired audio signal and an additive noise signal; and
applying noise suppression to the input audio signal to generate a noise-suppressed audio signal in a manner that is controlled by at least a parameter that specifies a degree of balance between distortion of the desired audio signal and unnaturalness of a residual noise signal included in the noise-suppressed audio signal.
2 . The method of claim 1 , further comprising:
determining the parameter that specifies the degree of balance between the distortion of the desired audio signal and the unnaturalness of the residual noise signal based at least in part on characteristics of the input audio signal.
3 . The method of claim 1 , wherein applying noise suppression to the input audio signal comprises:
passing a time domain representation of the input audio signal through a time domain filter having an impulse response that is controlled by at least the parameter that specifies the degree of balance between the distortion of the desired audio signal and the unnaturalness of the residual noise signal.
4 . The method of claim 3 , wherein passing the time domain representation of the input audio signal through the time domain filter comprises:
passing the time domain representation of the input audio signal through a time domain filter having an impulse response that is controlled by at least the parameter that specifies the degree of balance between the distortion of the desired audio signal and the unnaturalness of the residual noise signal and a noise attenuation factor.
5 . The method of claim 4 , further comprising:
identifying the noise attenuation factor; and
determining the degree of balance between the distortion of the desired audio signal and the unnaturalness of the residual noise signal based on the noise attenuation factor.
6 . The method of claim 3 , wherein passing the time domain representation of the input audio signal through the time domain filter comprises:
passing the time domain representation of the input audio signal through a time domain filter having an impulse response that is controlled by at least the parameter that specifies the degree of balance between the distortion of the desired audio signal and the unnaturalness of the residual noise signal and a noise shaping filter.
7 . The method of claim 1 , further comprising:
estimating statistics comprising correlation of the time domain representation of the input audio signal and correlation of a time domain representation of the additive noise signal; and
wherein passing the time domain representation of the input audio signal through the time domain filter comprises passing the time domain representation of the input audio signal through a time domain filter having an impulse response that is a function of at least the parameter that specifies the degree of balance between the distortion of the desired audio signal and the unnaturalness of the residual noise signal and the estimated statistics.
8 . The method of claim 1 , wherein applying noise suppression to the input audio signal comprises:
multiplying a frequency domain representation of the input audio signal by a frequency domain gain function that is controlled by at least the parameter that specifies the degree of balance between the distortion of the desired audio signal and the unnaturalness of the residual noise signal.
9 . The method of claim 8 , wherein multiplying the frequency domain representation of the input audio signal by the frequency domain gain function comprises multiplying the frequency domain representation of the input audio signal by a frequency domain gain function that is controlled by a single parameter that specifies the degree of balance between the distortion of the desired audio signal and the unnaturalness of the residual noise signal for all of a plurality of frequency sub-bands.
10 . The method of claim 8 , wherein multiplying the frequency domain representation of the input audio signal by the frequency domain gain function comprises multiplying the frequency domain representation of the input audio signal by a frequency domain gain function that is controlled by a plurality of parameters that specify the degree of balance between the distortion of the desired audio signal and the unnaturalness of the residual noise signal for each of a plurality of frequency sub-bands.
11 . The method of claim 8 , wherein multiplying the frequency domain representation of the input audio signal by the frequency domain gain function comprises:
multiplying the frequency domain representation of the input audio signal by a frequency domain gain function that is controlled by at least the parameter that specifies the degree of balance between the distortion of the desired audio signal and the unnaturalness of the residual noise signal and a frequency-dependent noise attenuation factor.
12 . The method of claim 8 , further comprising:
estimating statistics comprising power spectra associated with the input audio signal and power spectra associated with the additive noise signal;
wherein multiplying the frequency domain representation of the input audio signal by the frequency domain gain function comprises multiplying the frequency domain representation of the input audio signal by a frequency domain gain function that is a function of at least the parameter that specifies the degree of balance between the distortion of the desired audio signal and the unnaturalness of the residual noise signal and the estimated statistics.
13 . A method, comprising:
receiving a first input audio signal that comprises a first desired audio signal and a first additive noise signal;
receiving a second input audio signal that comprises a second desired audio signal and a second additive noise signal;
processing the first input audio signal to generate a first processed audio signal in a manner that is controlled by at least a parameter that specifies a degree of balance between distortion of the first desired audio signal and unnaturalness of a residual noise signal included in a noise-suppressed audio signal;
processing the second input audio signal to generate a second processed audio signal in a manner that is controlled by at least the parameter that specifies the degree of balance between distortion of the first desired audio signal and unnaturalness of the residual noise signal; and
combining at least the first processed audio signal and the second processed audio signal to produce the noise-suppressed audio signal.
14 . The method of claim 13 , further comprising:
determining the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal based at least in part on characteristics of the first input audio signal and/or characteristics of the second input audio signal.
15 . The method of claim 13 , wherein processing the first input audio signal comprises passing a time domain representation of the first input audio signal through a first time domain filter having an impulse response that is controlled by at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal;
wherein processing the second input audio signal comprises passing a time domain representation of the second input audio signal through a second time domain filter having an impulse response that is controlled by at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal; and
wherein combining at least the first processed audio signal and the second processed audio signal comprises adding the output of the first time domain filter to the output of the second time domain filter.
16 . The method of claim 15 , wherein passing the time domain representation of the first input audio signal through the first time domain filter comprises passing the time domain representation of the first input audio signal through a first time domain filter having an impulse response that is controlled by at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal and a noise attenuation factor; and
wherein passing the time domain representation of the second input audio signal through the second time domain filter comprises passing the time domain representation of the second input audio signal through a second time domain filter having an impulse response that is controlled by at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal and the noise attenuation factor.
17 . The method of claim 16 , further comprising:
identifying the noise attenuation factor; and
determining the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal based on the noise attenuation factor.
18 . The method of claim 15 , wherein passing the time domain representation of the first input audio signal through the first time domain filter comprises passing the time domain representation of the first input audio signal through a first time domain filter having an impulse response that is controlled by at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal and a noise shaping filter; and
wherein passing the time domain representation of the second input audio signal through the second time domain filter comprises passing the time domain representation of the second input audio signal through a second time domain filter having an impulse response that is controlled by at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal and the noise shaping filter.
19 . The method of claim 15 , further comprising:
estimating statistics that include correlation of the time domain representation of the first input audio signal, correlation of a time domain representation of the first additive noise signal, correlation of the time domain representation of the second input audio signal, correlation of a time domain representation of the second additive noise signal, a cross-correlation between the time domain representation of the first input audio signal and the time domain representation of the second input audio signal, and a cross-correlation of the time domain representation of the first additive noise signal and the time domain representation of the second additive noise signal; and
wherein passing the time domain representation of the first input audio signal through the first time domain filter comprises passing the time domain representation of the first input audio signal through a first time domain filter having an impulse response that is a function of at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal and at least some of the statistics; and
wherein passing the time domain representation of the second input audio signal through the second time domain filter comprises passing the time domain representation of the second input audio signal through a second time domain filter having an impulse response that is a function of at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal and at least some of the statistics.
20 . The method of claim 13 , wherein processing the first input audio signal comprises multiplying a frequency domain representation of the first input audio signal by a first frequency domain gain function that is controlled by at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal to generate a first product;
wherein processing the second input audio signal comprises multiplying a frequency domain representation of the second input audio signal by a second frequency domain gain function that is controlled by at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal to generate a second product; and
wherein combining at least the first processed audio signal and the second processed audio signal comprises adding the first product to the second product.
21 . The method of claim 20 , wherein
multiplying the frequency domain representation of the first input audio signal by the first frequency domain gain function comprises multiplying the frequency domain representation of the first input audio signal by a first frequency domain gain function that is controlled by a single parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal for all of a plurality of frequency sub-bands; and
multiplying the frequency domain representation of the second input audio signal by the second frequency domain gain function comprises multiplying the frequency domain representation of the second input audio signal by a second frequency domain gain function that is controlled by the single parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal for all of the plurality of frequency sub-bands.
22 . The method of claim 20 , wherein
multiplying the frequency domain representation of the first input audio signal by the first frequency domain gain function comprises multiplying the frequency domain representation of the first input audio signal by a first frequency domain gain function that is controlled by a plurality of parameters that specify the degree of balance between the distortion of the desired audio signal and the unnaturalness of the residual noise signal for each of a plurality of frequency sub-bands; and
multiplying the frequency domain representation of the second input audio signal by the second frequency domain gain function comprises multiplying the frequency domain representation of the second input audio signal by a second frequency domain gain function that is controlled by the plurality of parameters that specify the degree of balance between the distortion of the desired audio signal and the unnaturalness of the residual noise signal for each of the plurality of frequency sub-bands.
23 . The method of claim 20 , wherein multiplying the frequency domain representation of the first input audio signal by the first frequency domain gain function comprises multiplying the frequency domain representation of the first input audio signal by a first frequency domain gain function that is controlled by at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal and a frequency-dependent noise attenuation factor; and
wherein multiplying the frequency domain representation of the second input audio signal by the second frequency domain gain function comprises multiplying the frequency domain representation of the second input audio signal by a second frequency domain gain function that is controlled by at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal and the frequency-dependent noise attenuation factor.
24 . The method of claim 20 , further comprising:
estimating statistics comprising power spectra associated with the first input audio signal, power spectra associated with the second input audio signal, power spectra associated with the first additive noise signal, power spectra associated with the second additive noise signal, cross-power-spectra associated with the first and second input audio signals, and cross-power-spectra associated with the first and second additive noise signals;
wherein multiplying the frequency domain representation of the first input audio signal by the first frequency domain gain function comprises multiplying the frequency domain representation of the first input audio signal by a first frequency domain gain function that is a function of at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal and at least some of the statistics; and
wherein multiplying the frequency domain representation of the second input audio signal by the second frequency domain gain function comprises multiplying the frequency domain representation of the second input audio signal by a second frequency domain gain function that is a function of at least the parameter that specifies the degree of balance between the distortion of the first desired audio signal and the unnaturalness of the residual noise signal and at least some of the statistics.
25 . A method for applying noise suppression to an input audio signal, comprising:
receiving a plurality of sub-band signals obtained by applying a frequency conversion process to a time domain representation of the input audio signal; and
applying noise suppression to each of the sub-band signals by passing each of the sub-band signals through a corresponding time direction filter.
26 . The method of claim 25 , further comprising:
applying a time domain conversion process to the outputs of each of the corresponding time direction filters to generate a time domain representation of a noise-suppressed version of the input audio signal.
27 . The method of claim 25 , wherein receiving the plurality of sub-band signals comprises receiving the plurality of sub-band signals from a sub-band acoustic echo cancellation module.
28 . The method of claim 25 , wherein each sub-band signal comprises a desired audio signal and a noise signal; and
wherein passing each of the sub-band signals through a corresponding time direction filter comprises passing each of the sub-band signals through a time direction filter having a response that is controlled by at least a parameter that specifies a degree of balance between distortion of the desired audio signal included in the sub-band signal and unnaturalness of a residual noise signal included in a noise-suppressed version of the sub-band signal.
29 . The method of claim 28 , further comprising:
determining the parameter that specifies the degree of balance between the distortion of the desired audio signal included in the sub-band signal and the unnaturalness of the residual noise signal included in the noise-suppressed version of the sub-band signal for each sub-band based at least in part on characteristics of the input audio signal.
30 . The method of claim 28 , wherein passing each of the sub-band signals through a corresponding time direction filter comprises:
passing each of the sub-band signals through a corresponding time direction filter having a response that is controlled by at least a parameter that specifies the degree of balance between the distortion of the desired audio signal included in the sub-band signal and the unnaturalness of the residual noise signal included in the noise-suppressed version of the sub-band signal and a noise attenuation factor.
31 . The method of claim 30 , further comprising, for each sub-band:
identifying the noise attenuation factor; and
determining the degree of balance between the distortion of the desired audio signal included in the sub-band signal and the unnaturalness of the residual noise signal included in the noise-suppressed version of the sub-band signal based on the noise attenuation factor.
32 . The method of claim 28 , wherein passing each of the sub-band signals through a corresponding time direction filter comprises:
passing each of the sub-band signals through a corresponding time direction filter having a response that is controlled by at least a parameter that specifies the degree of balance between the distortion of the desired audio signal included in the sub-band signal and the unnaturalness of the residual noise signal included in the noise-suppressed version of the sub-band signal and a noise shaping filter.
33 . A method for performing noise suppression, comprising:
receiving a plurality of first sub-band signals obtained by applying a frequency conversion process to a time domain representation of a first input audio signal;
receiving a plurality of second sub-band signals obtained by applying a frequency conversion process to a time domain representation of a second input audio signal;
passing each of the plurality of first sub-band signals through a corresponding one of a plurality of first time direction filters;
passing each of the plurality of second sub-band signals through a corresponding one of a plurality of second time direction filters; and
combining an output from each of the plurality of first time direction filters with an output from a corresponding one of the plurality of second time direction filters to generate a plurality of noise-suppressed sub-band signals.
34 . The method of claim 33 , further comprising:
applying a time domain conversion process to the plurality of noise-suppressed sub-band signals to generate a time domain representation of a noise-suppressed audio signal.
35 . The method of claim 33 ,
wherein passing each of the plurality of first sub-band signals through a corresponding one of a plurality of first time direction filters comprises passing each first sub-band signal through a corresponding first time direction filter for a given sub-band having a response that is controlled by at least a parameter that specifies a degree of balance between distortion of a desired audio signal included in the first sub-band signal for the given sub-band and unnaturalness of a residual noise signal present in a noise-suppressed sub-band signal generated for the given sub-band; and
wherein passing each of the plurality of second sub-band signals through a corresponding one of a plurality of second time direction filters comprises passing each second sub-band signal through a corresponding second time direction filter for a given sub-band having a response that is controlled by at least a parameter that specifies a degree of balance between distortion of a desired audio signal included in the first sub-band signal for the given sub-band and unnaturalness of a residual noise signal present in the noise-suppressed sub-band signal generated for the given sub-band.