Method for processing audio input data and a device thereof
A computer-implemented method for processing audio input data into processed audio data by using an audio device comprising a microphone, a processor device and a memory holding a plurality of neural networks is presented. The plurality of neural networks are associated with different room types, wherein each room type is associated with one or more reference room acoustic metrics. The method comprises obtaining, by the microphone, room response data, wherein the room response data is reflecting room acoustics of a room in which the audio device is placed, determining, by using the processor device, the one or more room acoustic metrics based on the room response data, and selecting, by using the processor device, a matching neural network among the plurality of neural networks by comparing the one or more room acoustic metrics with the one or more reference room acoustic metrics associated with the different room types associated with the plurality of neural networks.
1 . A computer-implemented method for processing audio input data into processed audio data by using an audio device comprising a microphone, an output transducer, a processor device and a memory holding a plurality of neural networks, wherein the plurality of neural networks are associated with different room types, wherein each room type is associated with one or more reference room acoustic metrics, said computer-implemented method comprising:
obtaining speech data originating from a far-end room via a data communication device, wherein the audio device is placed in a near-end room,
obtaining, by the microphone, room response data, wherein the room response data is reflecting room acoustics of a room in which the audio device is placed, wherein the room response data captured by the microphone is based on sound generated by the output transducer using the speech data originating from the far-end room and the reflecting room acoustics of the room,
applying an adaptive filter to the speech data to provide an estimate of the reflecting room acoustics of the speech data in the room to provide an estimated echo signal, wherein the adaptive filter is configured for linear echo cancellation of the estimated echo signal to provide an echo cancelled signal and to minimize the echo cancelled signal using a normalized least mean square method,
determining, by using the processor device, one or more room acoustic metrics based on the echo cancelled signal, wherein the one or more room acoustic metrics includes an impulse response and the determining includes determining a reverberation time based on the impulse response,
selecting, by using the processor device, a matching neural network among the plurality of neural networks by comparing at least the reverberation time with the one or more reference room acoustic metrics associated with the different room types associated with the plurality of neural networks,
processing the echo cancelled signal into the processed audio data by using the matching neural network, and
generating a clean speech signal for the near-end room by using the output transducer using the processed audio data.
2 . The computer-implemented method according to claim 1 , wherein the reverberation time is for a given frequency band or a set of frequency bands.
3 . The computer-implemented method according to claim 1 , wherein the plurality of neural networks comprise a generally trained neural network, and the generally trained neural network is selected as the matching neural network in case no matching neural network is found by comparing the one or more room acoustics metrics with the one or more reference room acoustic metrics.
4 . The computer-implemented method according to claim 1 , wherein the plurality of neural networks have been trained with different loss functions, wherein the different loss functions differ in terms of trade-offs between different distortion types.
5 . The computer-implemented method according to claim 1 , wherein the audio input data and the processed audio data are multi-channel audio data.
6 . The computer-implemented method according to claim 1 , further comprising:
transferring the processed audio data to a far-end device placed in the far-end room, wherein the far-end device is provided with an output transducer arranged to generate sound based on the processed audio data.
7 . A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processor devices of an audio device, the one or more programs comprising instructions for performing the method according to any one of the claim 1 .
8 . An audio device comprising:
a data communication device arranged to receive speech data from a far-end room,
a microphone configured to obtain room response data, wherein the room response data is reflecting room acoustics of a room in which the audio device is placed, wherein the room response data obtained by the microphone is based on sound generated by an output transducer of the audio device using the speech data originating from the far-end room and the reflecting room acoustics of the room,
a memory holding a plurality of neural networks, wherein the plurality of neural networks are associated with different room types, wherein each room type is associated with one or more reference room acoustic metrics,
a processor device configured to:
apply an impulse response estimator, the impulse response estimator comprising a linear echo canceller to estimate the impulse response of the speech data and the reflecting room acoustics of the room, the linear echo canceller including an adaptive filter to apply to the speech data to provide an estimate of the reflecting room acoustics of the speech data in the room to provide an estimated echo signal, wherein the adaptive filter is configured for linear echo cancellation of the estimated echo signal to provide an echo cancelled signal and to minimize the echo cancelled signal using a normalized least mean square method, and
determine one or more room acoustic metrics based on the echo cancelled signal, wherein the determining includes determining a reverberation time based on the impulse response, and to select a matching neural network among the plurality of neural networks by comparing at least the reverberation time with the one or more reference room acoustic metrics associated with the different room types associated with the plurality of neural networks, wherein
the processor device is arranged to process the echo cancelled signal into the processed audio data by using the matching neural network, and
the output transducer configured to output a clean speech signal for the near-end room using the processed audio data.
9 . The audio device according to claim 8 , wherein the reverberation time is for a given frequency band.
10 . The audio device according to claim 8 , wherein the plurality of neural networks comprise a generally trained neural network, and the generally trained neural network is selected as the matching neural network in case no matching neural network is found by comparing the one or more room acoustics metrics with the one or more reference room acoustic metrics.
11 . The audio device according to claim 8 , wherein the plurality of neural networks have been trained with different loss functions, wherein the different loss functions differ in in terms of trade-offs between different distortion types.