Audio system, audio device, and method for speaker extraction
A method for speech extraction in an audio device is disclosed. The method comprises obtaining a microphone input signal from one or more microphones including a first microphone. The method comprises applying an extraction model to the microphone input signal for provision of an output. The method comprises extracting a near speaker component in the microphone input signal according to the output of the extraction model being a machine-learning model for provision of a speaker output. The method comprises outputting the speaker output.
1 . A method for speech extraction in an audio device, the method comprising
obtaining input signal from one or more microphones including a first microphone;
performing a short-time Fourier transformation on the input signal to obtain a frequency domain representation of the input signal;
performing a power normalizing on the frequency domain representation of the input signal to obtain a power normalized input signal;
feeding the power normalized input signal into an extraction model,
wherein the extraction model is a machine-learning model; and
wherein the extraction model is trained based on clean speech signals and a set of reverberant speech signals;
determining one or more mask parameters based on output of the extraction model, wherein the one or more mask parameters comprise first mask parameters and second mask parameters;
extract a near speaker component based on the first mask parameters;
extract a far speaker component based on the second mask parameters;
wherein the far speaker component originates from a far speaker at a distance larger than approximately 30 cm from the first microphone, and the near speaker component originates from a near speaker at a distance less than approximately 30 cm from the first microphone.
2 . The method according to claim 1 , the method further comprising:
determining a near speaker signal based on the near speaker component, and outputting the near speaker signal as a speaker output.
3 . The method according to claim 2 , wherein the method further comprising performing inverse short-time Fourier transformation on the speaker output for provision of an electrical output signal.
4 . The method according to claim 1 , wherein the extracting of the near speaker component in the input signal comprises:
determining one or more mask parameters including a first mask parameter based on the output of the extraction model.
5 . The method according to claim 1 , wherein the machine-learning model is an off-line trained neural network.
6 . The method according to claim 1 , wherein the extraction model comprises deep neural network.
7 . The method according to claim 1 , wherein the obtaining of the input signal comprises performing short-time Fourier transformation on the input signal from one or more microphones for provision of the input signal.
8 . The method according to claim 1 , wherein the method further comprising extracting an ambient noise component in the input signal according to the output of the extraction model.
9 . The method according to claim 1 ,
wherein the obtaining of the input signal from one or more microphones including the first microphone comprises obtaining one or more of a first microphone input signal, a second microphone input signal, and a combined microphone input signal based on the first microphone input signal and second microphone input signal,
wherein the input signal is based on one or more of the first microphone input signal, the second microphone input signal, and the combined microphone input signal.
10 . An audio device comprising a processor, an interface, a memory, and one or more transducers, wherein the audio device is configured to perform the method according to claim 1 .