IP Library Granted Patent US 12,620,406
Granted Patent B2
US 12,620,406 · App. 18/466,711 · Granted May 5, 2026

System and method for speech enhancement in multichannel audio processing systems

Inventors: Stanislav Kruchinin (Vienna, AT); Dushyant Sharma (Mountain House, CA); Rong Gong (Vienna, AT)
Assignee: Microsoft Technology Licensing, LLC.
G10L21/16H04S7/30H04S2400/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,620,406
App. No.
18/466,711
Granted
May 5, 2026
Kind
B2
Abstract

A method, computer program product, and computing system for enhancement of audio signals received from a plurality of microphones. A multichannel audio signal is received from a plurality of microphones and is processed with a short-time discrete cosine transform (STDCT) to generate a real-valued spectral representation of the multichannel signal encoding both magnitude and phase information. Magnitude- and phase-dependent weights are generated, and an enhanced single-channel signal is produced based upon, at least in part, the spectral representation of the multichannel signal and the magnitude- and phase-dependent weights.

Claims (44)

1 . A method comprising:

receiving a multichannel audio signal from a multichannel audio frontend that includes a plurality of microphones, the multichannel audio signal including spatial information associated with the multichannel audio frontend;

processing the multichannel audio signal using a short-time discrete cosine transform (STDCT) to generate a real-valued spectral representation of the multichannel audio signal, the spatial information associated with the multichannel audio frontend being encoded within the real-valued spectral representation of the multichannel audio signal;

generating magnitude- and phase-dependent weights associated with the real-valued spectral representation of the multichannel audio signal and the spatial information encoded therein;

generating a single-channel representation of the multichannel signal based upon, at least in part, the spectral representation of the multichannel audio signal and the magnitude- and phase-dependent weights;

generating, by a neural network, direction of arrival (DOA) information based on the magnitude- and phase-dependent weights; and

transmitting the single-channel representation of the multichannel signal, along with metadata that includes the DOA information, to a cloud device, wherein the cloud device includes an automatic speech recognition (ASR) model configured to perform speech recognition on the single-channel representation of the multichannel signal, and wherein the cloud device is configured to use the DOA information for speaker localization or diarization to improve the speech recognition performance of the ASR model.

2 . The method of claim 1 , wherein the STDCT comprises a modified discrete cosine transform (MDCT).

3 . The method of claim 2 , wherein the MDCT comprises one of a floating point MDCT and an integer MDCT.

4 . The method of claim 2 , further comprising generating direction of arrival information for the multichannel signal based on, at least in part, the magnitude- and phase-dependent weights.

5 . The method of claim 1 , further comprising performing an inverse DCT on the single-channel representation to obtain an audio signal representation of the multichannel audio signal.

6 . The method of claim 1 , further comprising:

encoding the single-channel representation signal prior to transmission of the single-channel representation signal over a transmission channel.

7 . The method of claim 1 , wherein the single-channel representation of the multichannel signal is further based upon a direction of arrival (DOA) of the signal.

8 . The method of claim 1 , further comprising:

transmitting the single-channel representation of the multichannel signal to an automatic speech recognition (ASR) backend configured for single-channel speech recognition.

9 . A system comprising:

one or more processors; and

a memory storing programming instructions for execution by the one or more processors, the programming instructions, upon execution by the one or more processors, causing the system to perform the following operations:

receiving a multichannel audio signal from a multichannel audio frontend that includes a plurality of microphones, the multichannel audio signal including spatial information associated with the multichannel audio frontend;

processing the multichannel audio signal using a short-time discrete cosine transform (STDCT) to generate a real-valued spectral representation of the multichannel audio signal, the spatial information associated with the multichannel audio frontend being encoded within the real-valued spectral representation of the multichannel audio signal;

generating magnitude- and phase-dependent weights associated with the real-valued spectral representation of the multichannel audio signal and spatial information encoded therein;

generating a single-channel representation of the multichannel signal based upon, at least in part, the spectral representation of the multichannel audio signal and the magnitude- and phase-dependent weights;

generating, by a neural network, direction of arrival (DOA) information based on the magnitude- and phase-dependent weights; and

transmitting the single-channel representation of the multichannel signal, along with metadata that includes the DOA information, to a cloud device, wherein the cloud device includes an automatic speech recognition (ASR) model configured to perform speech recognition on the single-channel representation of the multichannel signal, and wherein the cloud device is configured to use the DOA information for speaker localization or diarization to improve the speech recognition performance of the ASR model.

10 . The system of claim 9 , wherein the STDCT comprises a modified discrete cosine transform (MDCT).

11 . The system of claim 10 , wherein the MDCT comprises one of a floating point MDCT and an integer MDCT.

12 . The system of claim 9 , wherein the programming instructions further cause the system to perform the following operation:

generating direction of arrival information for the multichannel signal based on, at least in part, the magnitude- and phase-dependent weights.

13 . The system of claim 9 , wherein the programming instructions further cause the system to perform the following operation:

performing an inverse STDCT on the single-channel representation to obtain an audio signal representation of the multichannel audio signal.

14 . The system of claim 9 , wherein the programming instructions further cause the system to perform the following operation:

encoding the single-channel representation signal prior to transmission of the single-channel representation signal over a transmission channel.

15 . A computer program product residing on a non-transitory computer readable medium having programming instructions stored thereon which, when executed by one or more processors of a system, cause the system to perform the following operations comprising:

receiving a multichannel audio signal from a multichannel audio frontend that includes a plurality of microphones, the multichannel audio signal including spatial information associated with the multichannel audio frontend;

processing the multichannel audio signal with a modified discrete cosine transform (MDCT) to generate a spectral representation of the multichannel audio signal, the spatial information associated with the multichannel audio frontend being encoded within the spectral representation of the multichannel audio signal;

generating magnitude- and phase-dependent weights associated with the spectral representation of the multichannel audio signal and spatial information encoded therein;

generating a single-channel representation of the multichannel signal based upon, at least in part, the spectral representation of the multichannel audio signal and the magnitude- and phase-dependent weights;

generating, by a neural network, direction of arrival (DOA) information based on the magnitude- and phase-dependent weights; and

transmitting the single-channel representation of the multichannel signal, along with metadata that includes the DOA information, to a cloud device, wherein the cloud device includes an automatic speech recognition (ASR) model configured to perform speech recognition on the single-channel representation of the multichannel signal, and wherein the cloud device is configured to use the DOA information for speaker localization or diarization to improve the speech recognition performance of the ASR model.

16 . The computer program product of claim 15 , wherein the MDCT comprises one of a floating point MDCT and an integer MDCT.

17 . The computer program product of claim 15 , further comprising generating direction of arrival information for the multichannel signal based on, at least in part, the magnitude- and phase-dependent weights.

18 . The computer program product of claim 15 , wherein the programming instructions further cause the system to perform the following operation:

performing an inverse MDCT on the single-channel representation to obtain an audio signal representation of the multichannel audio signal.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065530/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2023
From: KRUCHININ, STANISLAV; SHARMA, DUSHYANT; GONG, RONG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 064895/0611 →
Continuity (1)
Related Publication 20250087230A1 · Mar 13, 2025
References Cited (20)
US 20050165587A1 · Cheng · 2005 [cited by examiner]
US 20110103591A1 · Ojala · 2011 [cited by examiner]
US 20110311061A1 · Oshikiri · 2011 [cited by examiner]
US 20130272539A1 · Kim · 2013 [cited by examiner]
US 20150296319A1 · Shenoy · 2015 [cited by examiner]
US 20220399026A1 · Gong · 2022 [cited by examiner]
US 20230253000A1 · Kono · 2023 [cited by examiner]
US 20240048902A1 · Vilermo · 2024 [cited by examiner]
US 20240118363A1 · Yasuda · 2024 [cited by examiner]
Shuai, Chenhao, et al. “mdctGAN: Taming transformer-based GAN for speech super-resolution with modified DCT spectra.” arXiv preprint arXiv:2305.11104 (2023). (Year: 2023). [cited by examiner]
Zeghidour, Neil, et al. “Soundstream: An end-to-end neural audio codec.” IEEE/ACM Transactions on Audio, Speech, and Language Processing 30 (2021): 495-507. (Year: 2021). [cited by examiner]
Li, Qinglong, et al. “Real-time monaural speech enhancement with short-time discrete cosine transform.” arXiv preprint arXiv: 2102.04629 (2021). (Year: 2021). [cited by examiner]
Wang, Ye, Miikka Vilermo, and Leonid Yaroslavsky. “Energy compaction property of the MDCT in comparison with other transforms.” Audio Engineering Society Convention 109. Audio Engineering Society, 2000. (Year: 2000). [cited by examiner]
Geng, Chuang, and Lei Wang. “End-to-end speech enhancement based on discrete cosine transform.” 2020 IEEE International Conference on Artificial Intelligence and Computer Applications (ICAICA). IEEE, 2020. (Year: 2020). [cited by examiner]
Chen, Shuixian, Ruimin Hu, and Shuhua Zhang. “Estimating spatial cues for audio coding in MDCT domain.” 2009 IEEE International Conference on Multimedia and Expo. IEEE, 2009. (Year: 2009). [cited by examiner]
Zhang, Shuhua, Weibei Dou, and Huazhong Yang. “MDCT sinusoidal analysis for audio signals analysis and processing.” IEEE transactions on audio, speech, and language processing 21.7 (2013): 1403-1414. (Year: 2013). [cited by examiner]
Mariotte, Théo, et al. “Microphone array channel combination algorithms for overlapped speech detection.” Interspeech 2022 Human and Humanizing Speech Technology. 2022. (Year: 2022). [cited by examiner]
Koizumi, Yuma, et al. “End-to-end sound source enhancement using deep neural network in the modified discrete cosine transform domain.” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICAS… [cited by examiner]
Cheng, Corey. “Method for estimating magnitude and phase in the MDCT domain.” Audio Engineering Society Convention 116. Audio Engineering Society, 2004. (Year: 2004). [cited by examiner]
Jones, Daniel T., et al. “Microphone array coding preserving spatial information for cloud-based multichannel speech recognition.” 2022 30th European Signal Processing Conference (EUSIPCO). IEEE, 2022. (Year: 2022). [cited by examiner]