System and method for speech enhancement in multichannel audio processing systems
A method, computer program product, and computing system for enhancement of audio signals received from a plurality of microphones. A multichannel audio signal is received from a plurality of microphones and is processed with a short-time discrete cosine transform (STDCT) to generate a real-valued spectral representation of the multichannel signal encoding both magnitude and phase information. Magnitude- and phase-dependent weights are generated, and an enhanced single-channel signal is produced based upon, at least in part, the spectral representation of the multichannel signal and the magnitude- and phase-dependent weights.
1 . A method comprising:
receiving a multichannel audio signal from a multichannel audio frontend that includes a plurality of microphones, the multichannel audio signal including spatial information associated with the multichannel audio frontend;
processing the multichannel audio signal using a short-time discrete cosine transform (STDCT) to generate a real-valued spectral representation of the multichannel audio signal, the spatial information associated with the multichannel audio frontend being encoded within the real-valued spectral representation of the multichannel audio signal;
generating magnitude- and phase-dependent weights associated with the real-valued spectral representation of the multichannel audio signal and the spatial information encoded therein;
generating a single-channel representation of the multichannel signal based upon, at least in part, the spectral representation of the multichannel audio signal and the magnitude- and phase-dependent weights;
generating, by a neural network, direction of arrival (DOA) information based on the magnitude- and phase-dependent weights; and
transmitting the single-channel representation of the multichannel signal, along with metadata that includes the DOA information, to a cloud device, wherein the cloud device includes an automatic speech recognition (ASR) model configured to perform speech recognition on the single-channel representation of the multichannel signal, and wherein the cloud device is configured to use the DOA information for speaker localization or diarization to improve the speech recognition performance of the ASR model.
2 . The method of claim 1 , wherein the STDCT comprises a modified discrete cosine transform (MDCT).
3 . The method of claim 2 , wherein the MDCT comprises one of a floating point MDCT and an integer MDCT.
4 . The method of claim 2 , further comprising generating direction of arrival information for the multichannel signal based on, at least in part, the magnitude- and phase-dependent weights.
5 . The method of claim 1 , further comprising performing an inverse DCT on the single-channel representation to obtain an audio signal representation of the multichannel audio signal.
6 . The method of claim 1 , further comprising:
encoding the single-channel representation signal prior to transmission of the single-channel representation signal over a transmission channel.
7 . The method of claim 1 , wherein the single-channel representation of the multichannel signal is further based upon a direction of arrival (DOA) of the signal.
8 . The method of claim 1 , further comprising:
transmitting the single-channel representation of the multichannel signal to an automatic speech recognition (ASR) backend configured for single-channel speech recognition.
9 . A system comprising:
one or more processors; and
a memory storing programming instructions for execution by the one or more processors, the programming instructions, upon execution by the one or more processors, causing the system to perform the following operations:
receiving a multichannel audio signal from a multichannel audio frontend that includes a plurality of microphones, the multichannel audio signal including spatial information associated with the multichannel audio frontend;
processing the multichannel audio signal using a short-time discrete cosine transform (STDCT) to generate a real-valued spectral representation of the multichannel audio signal, the spatial information associated with the multichannel audio frontend being encoded within the real-valued spectral representation of the multichannel audio signal;
generating magnitude- and phase-dependent weights associated with the real-valued spectral representation of the multichannel audio signal and spatial information encoded therein;
generating a single-channel representation of the multichannel signal based upon, at least in part, the spectral representation of the multichannel audio signal and the magnitude- and phase-dependent weights;
generating, by a neural network, direction of arrival (DOA) information based on the magnitude- and phase-dependent weights; and
transmitting the single-channel representation of the multichannel signal, along with metadata that includes the DOA information, to a cloud device, wherein the cloud device includes an automatic speech recognition (ASR) model configured to perform speech recognition on the single-channel representation of the multichannel signal, and wherein the cloud device is configured to use the DOA information for speaker localization or diarization to improve the speech recognition performance of the ASR model.
10 . The system of claim 9 , wherein the STDCT comprises a modified discrete cosine transform (MDCT).
11 . The system of claim 10 , wherein the MDCT comprises one of a floating point MDCT and an integer MDCT.
12 . The system of claim 9 , wherein the programming instructions further cause the system to perform the following operation:
generating direction of arrival information for the multichannel signal based on, at least in part, the magnitude- and phase-dependent weights.
13 . The system of claim 9 , wherein the programming instructions further cause the system to perform the following operation:
performing an inverse STDCT on the single-channel representation to obtain an audio signal representation of the multichannel audio signal.
14 . The system of claim 9 , wherein the programming instructions further cause the system to perform the following operation:
encoding the single-channel representation signal prior to transmission of the single-channel representation signal over a transmission channel.
15 . A computer program product residing on a non-transitory computer readable medium having programming instructions stored thereon which, when executed by one or more processors of a system, cause the system to perform the following operations comprising:
receiving a multichannel audio signal from a multichannel audio frontend that includes a plurality of microphones, the multichannel audio signal including spatial information associated with the multichannel audio frontend;
processing the multichannel audio signal with a modified discrete cosine transform (MDCT) to generate a spectral representation of the multichannel audio signal, the spatial information associated with the multichannel audio frontend being encoded within the spectral representation of the multichannel audio signal;
generating magnitude- and phase-dependent weights associated with the spectral representation of the multichannel audio signal and spatial information encoded therein;
generating a single-channel representation of the multichannel signal based upon, at least in part, the spectral representation of the multichannel audio signal and the magnitude- and phase-dependent weights;
generating, by a neural network, direction of arrival (DOA) information based on the magnitude- and phase-dependent weights; and
transmitting the single-channel representation of the multichannel signal, along with metadata that includes the DOA information, to a cloud device, wherein the cloud device includes an automatic speech recognition (ASR) model configured to perform speech recognition on the single-channel representation of the multichannel signal, and wherein the cloud device is configured to use the DOA information for speaker localization or diarization to improve the speech recognition performance of the ASR model.
16 . The computer program product of claim 15 , wherein the MDCT comprises one of a floating point MDCT and an integer MDCT.
17 . The computer program product of claim 15 , further comprising generating direction of arrival information for the multichannel signal based on, at least in part, the magnitude- and phase-dependent weights.
18 . The computer program product of claim 15 , wherein the programming instructions further cause the system to perform the following operation:
performing an inverse MDCT on the single-channel representation to obtain an audio signal representation of the multichannel audio signal.