MODEL-BASED SIGNAL ENHANCEMENT SYSTEM
A signal processing system enhances a speech input signal. A signal reconstruction circuit receives the speech input signal and extracts a spectral envelope. The signal reconstruction circuit generates an excitation signal based on the input signal, and generates a reconstructed speech signal based on the extracted spectral envelope and an excitation signal. A combining circuit combines the noise reduced signal and the reconstructed speech signal. Signal reconstruction and signal combinations may be based on a signal-to-noise ratio of the speech signal or another input.
1 . A method for processing a speech input signal, comprising:
estimating an input-signal-to-noise ratio or a signal-to-noise ratio of the speech input signal;
generating an excitation signal corresponding to the speech input signal;
extracting a spectral envelope of the speech input signal;
generating a reconstructed speech signal based on the excitation signal and the extracted spectral envelope;
filtering the speech input signal with a noise reduction circuit to generate a noise reduced signal; and
combining the reconstructed speech signal and the noise reduced signal based on the input-signal-to-noise ratio or the signal-to-noise ratio to generate an enhanced speech output signal.
2 . The method according to claim 1 , further comprising:
calculating a weight corresponding to the reconstructed speech signal based on the input-signal-to-noise ratio or the signal-to-noise ratio to generate a weighted reconstructed speech signal;
calculating a weight corresponding to the noise reduced signal based on the input-signal-to-noise ratio or the signal-to-noise ratio to obtain a weighted noise reduced signal; and where
generating the enhanced speech output signal comprises combining the weighted reconstructed speech signal and the weighted noise reduced signal.
3 . The method according to claim 1 where estimating the input-signal-to-noise ratio or the signal-to-noise ratio further comprises:
estimating the short-time power density spectrum of noise corresponding to the speech input signal; and
determining a short-time spectrogram of the speech input signal.
4 . The method according to claim 3 , where estimating the short-time power density spectrum of the noise further comprises:
smoothing the short-time power density spectrum of the speech input signal in time to generate a first smoothed short-time power density spectrum;
smoothing the first smoothed short-time power density spectrum in a positive frequency direction to generate a second smoothed short-time power density spectrum;
smoothing the second smoothed short-time power density spectrum in a negative frequency direction to obtain a third smoothed short-time power density spectrum; and
determining a minimum of the third smoothed short-time power density spectrum for a discrete time index n and the estimated short-time power density spectrum of the noise for a discrete time index n−1.
5 . The method according to claim 1 , where the excitation signal is generated using an excitation codebook.
6 . The method according to claim 1 , where the reconstructed speech signal is based on an estimated spectral envelope derived from the extracted spectral envelope and a spectral envelope codebook.
7 . The method according to claim 6 , further comprising:
generating a prototype spectral envelope corresponding to the spectral envelope codebook, the prototype spectral envelope providing a best match to the extracted spectral envelope corresponding to portions of the speech input signal having an input-signal-to-noise ratio greater than a predetermined threshold; and
where the estimated spectral envelope further comprises:
the prototype spectral envelope best match; and
the extracted spectral envelope corresponding to portions of the speech input signal having an input-signal-to-noise ratio less than or equal to the predetermined threshold.
8 . The method according to claim 7 , further comprising generating the estimated spectral envelope as sub-bands based on a weighted sum of the extracted spectral envelope smoothed in frequency and the prototype spectral envelope best match.
9 . The method according to claim 8 , further comprising generating the excitation signal based on filtered excitation sub-band signals, where the filtered excitation sub-band signals are generated using a spread noise reduction filter.
10 . The method according to claim 1 , further comprising:
generating sub-band signals corresponding to the reconstructed speech signal;
generating sub-band signals corresponding to the noise reduced signal;
adapting phases of the sub-band signals corresponding to the reconstructed speech signal to phases of the sub-band signals corresponding to the noise reduced signal; and
where adapting the phases is based on the input-signal-to-noise ratio of the speech input signal.
11 . A computer-readable storage medium having processor executable instructions to process a speech input signal by performing the acts of:
estimating an input-signal-to-noise ratio or a signal-to-noise ratio of the speech input signal;
generating an excitation signal corresponding to the speech input signal;
extracting a spectral envelope of the speech input signal;
generating a reconstructed speech signal based on the excitation signal and the extracted spectral envelope;
filtering the speech input signal with a noise reduction circuit to generate a noise reduced signal; and
combining the reconstructed speech signal and the noise reduced signal based on the input-signal-to-noise ratio or the signal-to-noise ratio to generate an enhanced speech output signal.
12 . The computer-readable storage medium of claim 11 , further comprising processor executable instructions to cause a processor to perform the acts of:
calculating a weight corresponding to the reconstructed speech signal based on the input-signal-to-noise ratio or the signal-to-noise ratio to generate a weighted reconstructed speech signal;
calculating a weight corresponding to the noise reduced signal based on the input-signal-to-noise ratio or the signal-to-noise ratio to obtain a weighted noise reduced signal; and where
generating the enhanced speech output signal comprises combining the weighted reconstructed speech signal and the weighted noise reduced signal.
13 . The computer-readable storage medium of claim 11 , further comprising processor executable instructions to cause a processor to perform the acts of estimating the input-signal-to-noise ratio or the signal-to-noise ratio by:
estimating the short-time power density spectrum of noise corresponding to the speech input signal; and
determining a short-time spectrogram of the speech input signal.
14 . The computer-readable storage medium of claim 13 , further comprising processor executable instructions to cause a processor to perform the acts of estimating the short-time power density spectrum of the noise by
smoothing the short-time power density spectrum of the speech input signal in time to generate a first smoothed short-time power density spectrum;
smoothing the first smoothed short-time power density spectrum in a positive frequency direction to generate a second smoothed short-time power density spectrum;
smoothing the second smoothed short-time power density spectrum in a negative frequency direction to obtain a third smoothed short-time power density spectrum; and
determining a minimum of the third smoothed short-time power density spectrum for a discrete time index n and the estimated short-time power density spectrum of the noise for a discrete time index n−1.
15 . The computer-readable storage medium of claim 11 , further comprising processor executable instructions to cause a processor to perform the act of accessing an excitation codebook to generate the excitation signal.
16 . The computer-readable storage medium of claim 11 , further comprising processor executable instructions to cause a processor to perform the acts of generating the reconstructed speech signal based on an estimated spectral envelope derived from the extracted spectral envelope and a spectral envelope codebook.
17 . The computer-readable storage medium of claim 16 , further comprising processor executable instructions to cause a processor to perform the acts of:
generating a prototype spectral envelope corresponding to the spectral envelope codebook, the prototype spectral envelope providing a best match to the extracted spectral envelope corresponding to portions of the speech input signal having an input-signal-to-noise ratio greater than a predetermined threshold; and
where the estimated spectral envelope further comprises:
the prototype spectral envelope best match; and
the extracted spectral envelope corresponding to portions of the speech input signal having an input-signal-to-noise ratio less than or equal to the predetermined threshold.
18 . The computer-readable storage medium of claim 17 , further comprising processor executable instructions to cause a processor to perform the acts of generating the estimated spectral envelope as sub-bands based on a weighted sum of the extracted spectral envelope smoothed in frequency and the prototype spectral envelope best match.
19 . The computer-readable storage medium of claim 18 , further comprising processor executable instructions to cause a processor to perform the acts of generating the excitation signal based on filtered excitation sub-band signals, where the filtered excitation sub-band signals are generated using a spread noise reduction filter.
20 . The computer-readable storage medium of claim 11 , further comprising processor executable instructions to cause a processor to perform the acts of:
generating sub-band signals corresponding to the reconstructed speech signal;
generating sub-band signals corresponding to the noise reduced signal;
adapting phases of the sub-band signals corresponding to the reconstructed speech signal to phases of the sub-band signals corresponding to the noise reduced signal; and
where adapting the phases is based on the input-signal-to-noise ratio of the speech input signal.
21 . A signal processing system for enhancing a speech input signal, comprising:
a noise reduction circuit configured to receive the speech input signal and generate a noise reduced signal;
a signal reconstruction circuit configured to receive the speech input signal and extract a spectral envelope from the speech input signal, the signal reconstruction circuit further configured to
generate an excitation signal based on the speech input signal; and
generate a reconstructed speech signal based on the extracted spectral envelope and the excitation signal;
a signal combining circuit configured to combine the noise reduced signal and the reconstructed speech signal to generate an enhanced speech output signal; and
a control circuit configured to receive the speech input signal and control the signal reconstruction circuit and the signal combining circuit based on an input-signal-to-noise ratio or a signal-to-noise ratio of the speech input signal.
22 . The system according to claim 21 , further comprising:
at least one analysis filter bank configured to transform the speech input signal into speech input sub-band signals;
at least one synthesis filter bank configured to synthesize sub-band signals generated by the noise reduction circuit and/or the signal reconstruction circuit.
23 . The system according to claim 22 , where the signal reconstruction circuit further comprises:
an excitation codebook;
a spectral envelope codebook;
an excitation estimation circuit configured to generate the excitation signal based on the excitation codebook;
a spectral envelope estimation circuit configured to generate an estimated spectral envelope based on the spectral envelope codebook; and
where the signal reconstruction circuit generates the reconstructed speech signal based on the estimated spectral envelope and the excitation signal.
24 . The system according to claim 21 , where the control circuit determines the input-signal-to-noise ratio or the signal-to-noise ratio of the speech input signal, and deactivates the signal reconstruction circuit if the determined input-signal-to-noise ratio or the signal-to-noise ratio exceeds a predetermined threshold.