Unified post-filter for an audio filter system for vehicle
A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations. The operations include receiving, from a sensor array, an audio signal at a unified post-filter, converting, via a conversion function, the audio signal into a Short-Time Fourier Transform (STFT) domain, determining, based on the converted audio signal, a speech-presence probability, and determining, based on the speech-presence probability, a noise smoothing factor. The operations also include estimating, via the unified post-filter, a noise power spectral density based on the noise smoothing factor, estimating, via the unified post-filter, a steering vector of a desired source, and generating, via the unified post-filter, a directionality-based mask and a coherence-based mask. The operations further include generating, based on the directionality-based mask and the coherence-based mask, a residual echo spectrum estimation and setting one or more spectral shaping factors based on the residual echo spectrum estimation.
1 . A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising:
receiving, from a sensor array, an audio signal at a unified post-filter;
converting, via a conversion function, the audio signal into a Short-Time Fourier Transform (STFT) domain;
determining, based on the converted audio signal, a speech-presence probability;
determining, based on the speech-presence probability, a noise smoothing factor;
estimating, via the unified post-filter, a noise power spectral density based on the noise smoothing factor;
estimating, via the unified post-filter, a steering vector of a desired source;
generating, via the unified post-filter, a directionality-based mask and a coherence-based mask;
generating, based on the directionality-based mask and the coherence-based mask, a residual echo spectrum estimation; and
setting one or more spectral shaping factors based on the residual echo spectrum estimation.
2 . The method of claim 1 , wherein the audio signal includes the desired source, residual ambient noise, and a residual echo.
3 . The method of claim 1 , wherein estimating the noise power spectral density includes determining an active speaker probability.
4 . The method of claim 1 , wherein generating the directionality-based mask includes utilizing spatial information and distinguishing the desired source from a residual echo of the audio signal.
5 . The method of claim 1 , wherein generating the coherence-based mask includes masking an estimated echo of the audio signal.
6 . The method of claim 1 , wherein generating the residual echo spectrum estimation includes extracting a residual echo of the audio signal from an original echo of the audio signal.
7 . The method of claim 1 , further including implementing, via the unified post-filter, a parametric variant of a Wiener filter.
8 . An audio filter system for a vehicle, the audio filter system comprising:
data processing hardware; and
memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
receiving, from a sensor array, an audio signal at a unified post-filter;
converting, via a conversion function, the audio signal into a Short-Time Fourier Transform (STFT) domain;
determining, based on the converted audio signal, a speech-presence probability;
determining, based on the speech-presence probability, a noise smoothing factor;
estimating, via the unified post-filter, a noise power spectral density based on the noise smoothing factor;
estimating, via the unified post-filter, a steering vector of a desired source;
generating, via the unified post-filter, a directionality-based mask and a coherence-based mask;
generating, based on the directionality-based mask and the coherence-based mask, a residual echo spectrum estimation; and
setting one or more spectral shaping factors based on the residual echo spectrum estimation.
9 . The system of claim 8 , wherein the audio signal includes the desired source, residual ambient noise, and a residual echo.
10 . The system of claim 8 , wherein determining the noise power spectral density includes determining an active speaker probability.
11 . The system of claim 8 , wherein generating the directionality-based mask includes utilizing spatial information and distinguishing the desired source from a residual echo of the audio signal.
12 . The system of claim 8 , wherein generating the coherence-based mask includes masking an estimated echo of the audio signal.
13 . The system of claim 8 , wherein generating the residual echo spectrum estimation includes extracting a residual echo of the audio signal from an original echo of the audio signal.
14 . The system of claim 8 , further including implementing, via the unified post-filter, a parametric variant of a Wiener filter.
15 . An audio filter system for a vehicle, the audio filter system comprising:
data processing hardware; and
memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
receiving, from a sensor array, an audio signal at a unified post-filter;
converting, via a conversion function, the audio signal into a Short-Time Fourier Transform (STFT) domain;
determining, based on the converted audio signal, a speech-presence probability;
determining, based on the speech-presence probability, a noise smoothing factor;
estimating, via the unified post-filter, a noise power spectral density based on the noise smoothing factor;
estimating, via the unified post-filter, a steering vector of a desired source;
generating, via the unified post-filter, a directionality-based mask and a coherence-based mask;
generating, based on the directionality-based mask and the coherence-based mask, a residual echo spectrum estimation;
setting one or more spectral shaping factors based on the residual echo spectrum estimation; and
implementing, via the unified post-filter, a parametric variant of a Wiener filter.
16 . The system of claim 15 , wherein the audio signal includes the desired source, residual ambient noise, and a residual echo.
17 . The system of claim 15 , wherein determining the noise power spectral density includes determining an active speaker probability.
18 . The system of claim 15 , wherein generating the directionality-based mask includes utilizing spatial information and distinguishing the desired source from a residual echo of the audio signal.
19 . The system of claim 15 , wherein generating the coherence-based mask includes masking an estimated echo of the audio signal.
20 . The system of claim 15 , wherein generating the residual echo spectrum estimation includes extracting a residual echo of the audio signal from an original echo of the audio signal.