IP Library Granted Patent US 12,406,682
Granted Patent B2
US 12,406,682 · App. 17/512,506 · Granted Sep 2, 2025

Real-time low-complexity echo cancellation

Inventors: Zhaofeng Jia (Saratoga, CA); Yang Liu (Wetherby, GB); Qiyong Liu (Singapore, SG)
Assignee: Zoom Communications, Inc.
G10L21/0208G10L2021/02082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,682
App. No.
17/512,506
Granted
Sep 2, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media relate to a method for acoustic echo cancellation. The system inputs one or more signal representations into an acoustic echo cancellation network comprising one or more network blocks to generate a mask, each network block comprising one or more convolutional blocks, each convolutional block comprising one or more neural networks. The system combines the mask and a near-end audio signal representation to generate an echo-cancelled audio signal representation. The system generates an echo-cancelled audio signal based on the echo-cancelled audio signal representation.

Claims (40)

1. A computer-implemented method for acoustic echo cancellation, comprising:

generating a far-end audio signal representation, a near-end audio signal representation, and a linear output signal representation based on a far-end audio signal, a near-end audio signal, and a linear output signal, respectively;

inputting the far-end audio signal representation, the near-end audio signal representation, and the linear output signal representation into an AEC network comprising one or more network blocks to generate a mask, each network block comprising one or more convolutional blocks, each convolutional block comprising one or more neural networks;

combining the mask and the near-end audio signal representation to generate an echo-cancelled audio signal representation; and

generating an echo-cancelled audio signal based on the echo-cancelled audio signal representation.

2. The method of claim 1 , wherein the far-end audio signal representation, the near-end audio signal representation, and the linear output signal representation comprise STFTs of the far-end audio signal, the near-end audio signal, and the linear output signal, respectively.

3. The method of claim 2 , wherein the echo-cancelled audio signal is generated based on an inverse STFT of the echo-cancelled audio signal representation.

4. The method of claim 1 , wherein each network block comprises a series of convolutional blocks of increasing dilation, the output of each convolutional block in the series being input to the next convolutional block in the series.

5. The method of claim 4 , further comprising:

summing the outputs of one or more convolutional blocks in a network block and inputting the sum to a next network block.

6. The method of claim 5 , further comprising:

fusing the sum of the outputs of the one or more convolutional blocks in the network block with an embedding of the far-end audio signal representation, the near-end audio signal representation, and the linear output signal representation prior to inputting the sum to the next network block.

7. The method of claim 1 , wherein the far-end audio signal comprises a first speech audio signal, the near-end audio signal comprises a second speech audio signal combined with an echo of the far-end audio signal, and the linear output signal comprises output of a DSP AEC linear filter, and the AEC network is trained by minimizing a loss function based on the difference between the echo-cancelled audio signal and the second speech audio signal.

8. A non-transitory computer readable medium that stores executable program instructions that when executed by one or more computing devices configure the one or more computing devices to:

generate a far-end audio signal representation, a near-end audio signal representation, and a linear output signal representation based on a far-end audio signal, a near-end audio signal, and a linear output signal, respectively;

input the far-end audio signal representation, the near-end audio signal representation, and the linear output signal representation into an AEC network comprising one or more network blocks to generate a mask, each network block comprising one or more convolutional blocks, each convolutional block comprising one or more neural networks;

combine the mask and the near-end audio signal representation to generate an echo-cancelled audio signal representation; and

generate an echo-cancelled audio signal based on the echo-cancelled audio signal representation.

9. The non-transitory computer readable medium of claim 8 , wherein the far-end audio signal representation, the near-end audio signal representation, and the linear output signal representation comprise STFTs of the far-end audio signal, the near-end audio signal, and the linear output signal, respectively.

10. The non-transitory computer readable medium of claim 9 , wherein the echo-cancelled audio signal is generated based on an inverse STFT of the echo-cancelled audio signal representation.

11. The non-transitory computer readable medium of claim 8 , wherein each network block comprises a series of convolutional blocks of increasing dilation, the output of each convolutional block in the series being input to the next convolutional block in the series.

12. The non-transitory computer readable medium of claim 11 , further comprising executable program instructions that when executed by one or more computing devices configure the one or more computing devices to:

sum the outputs of one or more convolutional blocks in a network block and inputting the sum to a next network block.

13. The non-transitory computer readable medium of claim 12 , further comprising executable program instructions that when executed by one or more computing devices configure the one or more computing devices to:

fuse the sum of the outputs of the one or more convolutional blocks in the network block with an embedding of the far-end audio signal representation, the near-end audio signal representation, and the linear output signal representation prior to inputting the sum to the next network block.

14. The non-transitory computer readable medium of claim 8 , wherein the far-end audio signal comprises a first speech audio signal, the near-end audio signal comprises a second speech audio signal combined with an echo of the far-end audio signal, and the linear output signal comprises output of a DSP AEC linear filter, and the AEC network is trained by minimizing a loss function based on the difference between the echo-cancelled audio signal and the second speech audio signal.

15. An acoustic echo cancellation system comprising:

a non-transitory computer-readable medium; and

one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium, the processor-executable instructions configured to cause the one or more processors to:

generate a far-end audio signal representation, a near-end audio signal representation, and a linear output signal representation based on a far-end audio signal, a near-end audio signal, and a linear output signal, respectively;

input the far-end audio signal representation, the near-end audio signal representation, and the linear output signal representation into an acoustic echo cancellation (AEC) network comprising one or more network blocks to generate a mask, each network block comprising one or more convolutional blocks, each convolutional block comprising one or more neural networks;

combine the mask and the near-end audio signal representation to generate an echo-cancelled audio signal representation; and

generate an echo-cancelled audio signal based on the echo-cancelled audio signal representation.

16. The system of claim 15 , wherein the far-end audio signal representation, the near-end audio signal representation, and the linear output signal representation comprise Short-time Fourier Transforms (STFT) of the far-end audio signal, the near-end audio signal, and the linear output signal, respectively.

17. The system of claim 16 , wherein the echo-cancelled audio signal is generated based on an inverse STFT of the echo-cancelled audio signal representation.

18. The system of claim 15 , wherein each network block comprises a series of convolutional blocks of increasing dilation, the output of each convolutional block in the series being input to the next convolutional block in the series.

19. The system of claim 18 , wherein the one or more processors are further configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:

sum the outputs of one or more convolutional blocks in a network block and inputting the sum to a next network block.

20. The system of claim 19 , wherein the one or more processors are further configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:

fuse the sum of the outputs of the one or more convolutional blocks in the network block with an embedding of the far-end audio signal representation, the near-end audio signal representation, and the linear output signal representation prior to inputting the sum to the next network block.

Assignments (2)
CHANGE OF NAME Recorded Jun 3, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 071480/0463 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 27, 2021
From: JIA, ZHAOFENG; LIU, YANG; LIU, QIYONG
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 057938/0845 →
Priority Claims (1)
CN 202111122229.9 · Sep 24, 2021 · national
Continuity (1)
Related Publication 20230096565A1 · Mar 30, 2023
References Cited (15)
US 20130216057A1 · Thyssen · 2013 [cited by examiner]
US 20190222691A1 · Shah · 2019 [cited by examiner]
US 20220277721A1 · Zhang · 2022 [cited by examiner]
US 20230094630A1 · Zhang · 2023 [cited by examiner]
Luo, Yi, and Nima Mesgarani. “Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation.” IEEE/ACM transactions on audio, speech, and language processing 27, No. 8 (2019): 1256-1266. [cited by applicant]
Valin, Jean-Marc, Srikanth Tenneti, Karim Helwani, Umut Isik, and Arvindh Krishnaswamy. “Low-Complexity, Real- Time Joint Neural Echo Control and Speech Enhancement Based On Percepnet.” In ICASSP 2021-2021 IEEE Internat… [cited by applicant]
Chen, Hongsheng, Teng Xiang, Kai Chen, and Jing Lu. “Nonlinear Residual Echo Suppression Based on Multi-stream Conv-TasNet.” arXiv preprint arXiv:2005.07631 (2020). [cited by applicant]
Halimeh, Mhd Modar, and Walter Kellermann. “Efficient multichannel nonlinear acoustic echo cancellation based on a cooperative strategy.” In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal… [cited by applicant]
Pfeifenberger, Lukas, and Franz Pernkopf. “Nonlinear Residual Echo Suppression Using a Recurrent Neural Network.” In Interspeech, pp. 3950-3954. 2020. [cited by applicant]
Fan, Wenzhi, and Jing Lu. “Improving partition-block-based acoustic echo canceler in under-modeling scenarios.” arXiv preprint arXiv:2008.03944 (2020). [cited by applicant]
Valin, Jean-Marc. “A hybrid DSP/deep learning approach to real-time full-band speech enhancement.” In 2018 IEEE 20th international workshop on multimedia signal processing (MMSP), pp. 1-5. IEEE, 2018. [cited by applicant]
Halimeh, Mhd Modar, Thomas Haubner, Annika Briegleb, Alexander Schmidt, and Walter Kellermann. “Combining Adaptive Filtering And Complex-Valued Deep Postfiltering For Acoustic Echo Cancellation.” In ICASSP 2021-2021 EEE… [cited by applicant]
Zhang, Yi, Chengyun Deng, Shiqian Ma, Yongtao Sha, and Hui Song. “Deep Multi-task Network for Delay Estimation and Echo Cancellation.” arXiv preprint arXiv:2011.02109 (2020). [cited by applicant]
Bagheri, Saeed, and Daniele Giacobello. “Robust STFT Domain Multi-Channel Acoustic Echo Cancellation with Adaptive Decorrelation of the Reference Signals.” In ICASSP 2021-2021 IEEE International Conference on Acoustics,… [cited by applicant]
Kim, Eesung, Jae-Jin Jeon and Hyeji Seo. “U-Convolution Based Residual Echo Suppression with Multiple Encoders.” ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2021):… [cited by applicant]