IP Library Granted Patent US 12,670,920
Granted Patent B2
US 12,670,920 · App. 18/172,017 · Granted Jun 30, 2026

Joint acoustic echo cancellation (AEC) and personalized noise suppression (PNS)

Inventors: Sefik Emre Eskimez (Bellevue, WA); Takuya Yoshioka (Bellevue, WA); Huaming Wang (Clyde Hill, WA); Alex Chenzhi Ju (Seattle, WA); Min Tang (Redmond, WA); Tanel Pärnamaa (Tallinn, EE)
Assignee: Microsoft Technology Licensing, LLC
G10L21/0232G06N3/0442G10L17/02G10L17/04G10L17/06G10L17/18G10L2021/02082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,920
App. No.
18/172,017
Granted
Jun 30, 2026
Kind
B2
Abstract

A data processing system implements receiving a far-end signal associated with a first computing device participating in an online communication session and receiving a near-end signal associated with a second computing device participating in the online communication session. The near-end signal includes speech of a target speaker, a first interfering speaker, and an echo signal. The system further implements providing the far-end signal, the near-end signal, and an indication of the target speaker as an input to a machine learning model. The machine learning model trained to analyze the far-end signal and the near-end signal to perform personalized noise suppression (PNS) to remove speech from one or more interfering speakers and acoustic echo cancellation (AEC) to remove echoes. The model is trained to output an audio signal comprising speech of the target speaker. The system obtains the audio signal comprising the speech of the target speaker from the model.

Claims (60)

1 . A data processing system comprising:

a processor; and

a machine-readable medium storing executable instructions that, when executed, cause the processor to perform operations comprising:

receiving a far-end signal associated with a first computing device participating in an online communication session;

encoding the far-end signal using a first learnable encoder to generate first embeddings representing first features of the far-end signal;

receiving a near-end signal associated with a second computing device participating in the online communication session, the near-end signal including speech of a target speaker, one or more interfering speakers, and an echo signal;

encoding the near-end signal using a second learnable encoder to generate second embeddings representing second features of the near-end signal;

providing the first embeddings representing features of the far-end signal, the second embeddings representing features of the near-end signal, and an indication of the target speaker as an input to a machine learning model comprising:

an alignment block, the alignment block configured to use attention to align the first features of the far-end signal included in the first embeddings and the second features of the near-end signal included in the second embeddings, the alignment block being further configured to output attention weights;

a first concatenation block configured to concatenate the first embeddings, the second embeddings, and the attention weights into a first concatenated input,

a first group of Long Short-Term Memory (LSTM) blocks trained to perform acoustic echo cancellation (AEC) on the first concatenated input to remove an echo and to output third features,

a second concatenation block configured to concatenate the third features output by the first group of LSTM blocks and the attention weights output by the alignment block into a second concatenated input, and

a second group of LSTM blocks trained to perform personalized noise suppression (PNS) on the second concatenated input to remove speech from one or more interfering speakers and to remove a residual echo output by the first group of LSTM blocks; and

obtaining an audio signal comprising the speech of the target speaker from the machine learning model.

2 . The data processing system of claim 1 , wherein the indication of the target speaker comprises a target speaker embedding vector representing speech characteristics of the target speaker.

3 . The data processing system of claim 2 , wherein the machine-readable medium includes instructions configured to cause the processor to perform operations of:

generating the target speaker embedding vector by capturing audio content comprising speech of the target speaker and extracting features from the audio content.

4 . The data processing system of claim 2 , wherein the second group of LSTM blocks receives the target speaker embedding vector as an input.

5 . The data processing system of claim 1 , wherein the machine-readable medium includes instructions configured to cause the processor to perform operations of sending the audio signal to the first computing device.

6 . A method implemented in a data processing system for processing audio signals, the method comprising:

receiving a far-end signal associated with a first computing device participating in an online communication session;

encoding the far-end signal using a first learnable encoder to generate first embeddings representing first features of the far-end signal;

receiving a near-end signal associated with a second computing device participating in the online communication session, the near-end signal including speech of a target speaker, one or more first interfering speakers, and an echo signal;

encoding the near-end signal using a second learnable encoder to generate second embeddings representing second features of the near-end signal;

analyzing the first embeddings representing features of the far-end signal, the second embeddings representing features of the near-end signal, and an indication of the target speaker as an input to a machine learning model by:

aligning the first features of the far-end signal included in the first embeddings and the second features of the near-end signal included in the second embeddings using an alignment block configured to use attention to align the first features and the second features;

obtaining attention weights output by the alignment block in response to aligning the first features and the second features;

concatenating the first embeddings, the second embeddings, and the attention weights into a first concatenated input using a first concatenation block;

performing acoustic echo cancellation (AEC) on the first concatenated input using a first group of Long Short-Term Memory (LSTM) blocks to remove an echo and to output third features,

concatenating the third features output by the first group of LSTM blocks and the attention weights output by the alignment block into a second concatenated input using a second concatenation block,

performing personalized noise suppression (PNS) on the second concatenated input using a second group of LSTM blocks to remove speech from the one or more interfering speakers and to remove a residual echo output by the first group of LSTM blocks, and

outputting an audio signal comprising speech of the target speaker; and

obtaining the audio signal comprising the speech of the target speaker from the machine learning model.

7 . The method of claim 6 , wherein the indication of the target speaker comprises a speaker embedding vector representing speech characteristics of the target speaker.

8 . The method of claim 7 , further comprising generating the speaker embedding vector by capturing audio content comprising speech of the target speaker and extracting features from the audio content.

9 . The method of claim 6 , further comprising performing AEC on features extracted from the near-end signal to remove echoes from the near-end signal using the machine learning model.

10 . The method of claim 9 , further comprising aligning features of the near-end signal and features of the far-end signal using attention using an alignment block.

11 . The method of claim 10 , further comprising providing alignment information output by the alignment block as an input to the first group of LSTM blocks.

12 . A data processing system comprising:

a processor; and

a machine-readable medium storing executable instructions that, when executed, cause the processor to perform operations comprising:

training a machine learning model using a first batch of training data to train the machine learning model to perform acoustic echo cancellation (AEC) to remove echoes from an input audio signal;

training the machine learning model using a second batch of training data to train the machine learning model to perform personalized noise suppression (PNS) to extract speech associated with a target speaker from the input audio signal;

training the machine learning model using a third batch of training data to train the machine learning model to perform both AEC and PNS on the input audio signal; and

analyzing audio signals associated with a communication session using the machine learning model to remove echoes and to extract the speech of a target speaker participating in the communication session by:

receiving a far-end signal associated with a first computing device participating in an online communication session;

encoding the far-end signal using a first learnable encoder to generate first embeddings representing first features of the far-end signal;

receiving a near-end signal associated with a second computing device participating in the online communication session, the near-end signal including speech of a target speaker, one or more interfering speakers, and an echo signal;

encoding the near-end signal using a second learnable encoder to generate second embeddings representing second features of the near-end signal;

analyzing the first embeddings representing features of the far-end signal, the second embeddings representing features of the near-end signal, and an indication of the target speaker as an input to a machine learning model by:

aligning the first features of the far-end signal included in the first embeddings and the second features of the near-end signal included in the second embeddings using an alignment block configured to use attention to align the first features and the second features;

obtaining attention weights output by the alignment block in response to aligning the first features and the second features;

concatenating the first embeddings, the second embeddings, and the attention weights into a first concatenated input using a first concatenation block;

performing acoustic echo cancellation (AEC) on the first concatenated input using a first group of Long Short-Term Memory (LSTM) blocks to remove the echo signal and to output third features,

concatenating the third features output by the first group of LSTM blocks and the attention weights output by the alignment block into a second concatenated input using a second concatenation block, and

performing personalized noise suppression (PNS) on the second concatenated input using a second group of LSTM blocks to remove speech from the one or more interfering speakers and to remove a residual echo output by the first group of LSTM blocks, and

outputting an audio signal comprising speech of the target speaker.

13 . The data processing system of claim 12 , wherein the first batch of training data includes speech of the target speaker, noise data, and echo data.

14 . The data processing system of claim 13 , wherein the second batch of training data includes the speech of the target speaker, the speech of the one or more interfering speakers who are different than the target speaker, and the noise data.

15 . The data processing system of claim 14 , wherein the second batch of training data further includes echo data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2023
From: ESKIMEZ, SEFIK EMRE; YOSHIOKA, TAKUYA; WANG, HUAMING; JU, ALEX CHENZHI; TANG, MIN; PÄRNAMAA, TANEL
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 062757/0195 →
Continuity (2)
Provisional Application 63417182 · Oct 18, 2022
Related Publication 20240135949A1 · Apr 25, 2024
References Cited (120)
US 11393487B2 · Fazeli et al. · 2022 [cited by applicant]
US 20220180886A1 · Weng · 2022 [cited by applicant]
US 20230403505A1 · Yu · 2023 [cited by examiner]
CN 112259112A · 2021 [cited by applicant]
CN 112687288A · 2021 [cited by applicant]
G. Mittag and S. Möller, “Full-Reference Speech Quality Estimation with Attentional Siamese Neural Networks,” ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona… [cited by examiner]
Mendes, Andre, Julian Togelius, and Leandro dos Santos Coelho. “Multi-stage transfer learning with an application to selection process.” ECAI 2020. IOS Press, 2020. 1770-1777 (Year: 2020). [cited by examiner]
Eskimez, et al., “Real-Time Joint Personalized Speech Enhancement and Acoustic Echo Cancellation”, Arxiv, May 25, 2023, 5 pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US23/033316, mailed on Jan. 8, 2024, 16 pages. [cited by applicant]
Zhang et al., “Personalized Acoustic Echo Cancellation for Full-duplex Communications”, Arxiv, Jun. 30, 2022, 5 pages. [cited by applicant]
Indenbom, et al., “Deep model with built-in self-attention alignment for acoustic echo cancellation”, Arxiv, Aug. 24, 2022, 5 pages. [cited by applicant]
International Preliminary Report On Patentability received for PCT Application No. PCT/US23/033316, mailed on May 1, 2025, 10 pages. [cited by applicant]
Meng, et al., “Neural Echo: A Self-Attentive Recurrent Neural Network For Unified Acoustic Echo Suppression And Speech Enhancement”, Arxiv, May 20, 2022, 5 pages. [cited by applicant]
Nils L, et al., “Acoustic Echo Cancellation with the Dual-Signal Transformation LSTM Network”, IEEE, Jun. 6, 2021, pp. 7138-7142. [cited by applicant]
“AEC-Challenge”, Retrieved from: https://github.com/microsoft/AEC-Challenge, Retrieved on Oct. 5, 2022, 4 Page. [cited by applicant]
“ITU-T P.831 : Subjective Performance Evaluation of Network Echo Cancellers”, Retrieved from: https://www.itu.int/ITU-T/recommendations/rec.aspx?rec=4537&lang=en, Dec. 3, 1998, 33 Pages. [cited by applicant]
“ITU-T P.832 : Subjective Performance Evaluation of Hands-Free Terminals”, Retrieved from: https://www.itu.int/ITU-T/recommendations/rec.aspx?rec=5086&lang=en, May 18, 2000, 29 Pages. [cited by applicant]
Knaderi, et al., “Microsoft / P.808”, Retrieved from: https://github.com/microsoft/P.808, Retrieved on Oct. 5, 2022 5 Pages. [cited by applicant]
“Microsoft / AEC-Challenge”, Retrieved from: https://github.com/microsoft/AEC-Challenge/tree/main/baseline/icassp2022, Retrieved on Oct. 5, 2022, 1 Page. [cited by applicant]
“Perceptual Evaluation of Speech Quality (PESQ): An Objective Method for End-to-End Speech Quality Assessment of Narrow-band Telephone Networks and Speech Codecs”, In ITU-T Recommendation P.862, Feb. 23, 2001, 30 Pages. [cited by applicant]
Hershey, et al., “Deep Clustering: Discriminative Embeddings for Segmentation and Separation”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, May 19, 2016, pp. 31-35. [cited by applicant]
Karjalainen, et al., “Estimation of Modal Decay Parameters from Noisy Response Measurements”, In 110th Convention of the Audio Engineering Society, May 12, 2001, pp. 867-878. [cited by applicant]
Avila, et al., “Non-intrusive Speech Quality Assessment using Neural Networks”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, May 12, 2019, pp. 631-635. [cited by applicant]
Benesty, et al., “Advances in Network and Acoustic Echo Cancellation”, In Proceedings of Digital Signal Processing, Jan. 2021, 232 Pages. [cited by applicant]
Farhang-Boroujeny, et al., “Adaptive Filters: Theory and Applications”, In Publication of A John Wiley & Sons, Ltd., Apr. 2, 2013, 802 Pages. [cited by applicant]
Valentini-Botinhao, et al., “Speech Enhancement for a Noise-Robust Text-to-Speech Synthesis System Using Deep Recurrent Neural Networks”, In Proceedings of INTERSPEECH, Sep. 2016, 5 Pages. [cited by applicant]
Braun, et al., “Towards Efficient Models for Real-Time Deep Noise Suppression”, In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Jun. 6, 2021, pp. 656-660. [cited by applicant]
Bucila, et al., “Model Compression”, In Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Aug. 20, 2006, pp. 535-541. [cited by applicant]
Casebeer, et al., “Auto-DSP: Learning to Optimize Acoustic Echo Cancellers”, In Proceedings of IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, Oct. 17, 2021, pp. 291-295. [cited by applicant]
Chang, et al., “DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden-unit BERT”, In Repository of arXiv:2110.01900v2, Oct. 6, 2021, 5 Pages. [cited by applicant]
Chen, et al., “Continuous Speech Separation with Conformer”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, May 13, 2021, pp. 5749-5753. [cited by applicant]
Chen, et al., “Deep Attractor Network for Single-microphone Speaker Separation”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Mar. 5, 2017, 11 Pages. [cited by applicant]
Chen, et al., “Ultra Fast Speech Separation Model with Teacher Student Learning”, In Proceedings of INTERSPEECH, Aug. 30, 2021, pp. 3026-3030. [cited by applicant]
Choi, et al., “Phase-Aware Speech Enhancement with Deep Complex U-Net”, In Proceedings of International Conference on Learning Representations, Mar. 7, 2019, 20 Pages. [cited by applicant]
Yamagishi, et al., “CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit”, Retrieved From: https://datashare.ed.ac.uk/handle/10283/3443, Nov. 13, 2019, 2 Pages. [cited by applicant]
Chung, et al., “Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling”, In Repository of arXiv:1412.3555, Dec. 11, 2014, 9 Pages. [cited by applicant]
Clevert, et al., “Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)”, In Repository of arXiv:1511.07289v2, Dec. 3, 2015, 14 Pages. [cited by applicant]
Cui, et al., “Multi-Scale Refinement Network Based Acoustic Echo Cancellation”, In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, May 23, 2022, pp. 9132-9136. [cited by applicant]
Cutler, et al., “Crowdsourcing Approach for Subjective Evaluation of Echo Impairment”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Jun. 6, 2021, pp. 406-410. [cited by applicant]
Cutler, et al., “ICASSP 2022 Acoustic Echo Cancellation Challenge”, In Repository of arXiv:2202.13290v1, Feb. 27, 2022, 5 Pages. [cited by applicant]
Cutler, et al., “INTERSPEECH 2021 Acoustic Echo Cancellation Challenge”, In Proceedings of INTERSPEECH, Jun. 2021, 5 Pages. [cited by applicant]
Delcroix, et al., “Single Channel Target Speaker Extraction and Recognition with Speaker Beam”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Sep. 13, 2018, pp. 5554-5558. [cited by applicant]
Dubey, et al., “ICASSP 2022 Deep Noise Suppression Challenge”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 27, 2022, pp. 9271-9275. [cited by applicant]
Eneman, et al., “Iterated Partitioned Block Frequency-Domain Adaptive Filtering for Acoustic Echo Cancellation”, In IEEE Transactions on Speech and Audio Processing, vol. 11, Issue 2, Mar. 2003, pp. 143-158. [cited by applicant]
Enzner, et al., “Acoustic Echo Control”, In Proceedings of the Academic Press Library in Signal Processing, Jan. 1, 2014, pp. 807-877. [cited by applicant]
Ephrat, et al., “Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation”, In the Journal of ACM Transactions on Graphics, vol. 37, Issue 4, Aug. 2018, 11 Pages. [cited by applicant]
Eskimez, et al., “Human Listening and Live Captioning: Multi-Task Training for Speech Enhancement”, In Proceedings of INTERSPEECH, Aug. 30, 2021, pp. 2686-2690. [cited by applicant]
Eskimez, et al., “Personalized Speech Enhancement: New Models and Comprehensive Evaluation”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 27, 2022, pp. 356-360. [cited by applicant]
Fazel, et al., “CAD-AEC: Context-Aware Deep Acoustic Echo Cancellation”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, May 4, 2020, pp. 6919-6923. [cited by applicant]
Fonseca, et al., “Freesound Datasets: A Platform for the Creation of Open Audio Datasets”, In Proceedings of the 18th ISMIR Conference, Oct. 23, 2017, pp. 486-493. [cited by applicant]
Furlanello, et al., “Born Again Neural Networks”, In Proceedings of the International Conference on Machine Learning, Jul. 3, 2018, 10 Pages. [cited by applicant]
Gamper, et al., “Blind Reverberation Time Estimation Using a Convolutional Neural Network”, In Proceedings of 16th International Workshop on Acoustic Signal Enhancement, Sep. 2018, pp. 136-140. [cited by applicant]
Gamper, et al., “Intrusive and Non-Intrusive Perceptual Speech Quality Assessment Using a Convolutional Neural Network”, In Proceedings of IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, Oct. … [cited by applicant]
Garofolo, et al., “DARPA TIMIT: Acoustic-Phonetic Continuous Speech Corpus CD-ROM”, In NIST Speech Disc, Feb. 1993, 94 Pages. [cited by applicant]
Gemmeke, et al., “Audio Set: An Ontology and Human-labeled Dataset for Audio Events”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Jun. 19, 2017, pp. 776-780. [cited by applicant]
Giri, et al., “Personalized PercepNet: Real-time, Low-complexity Target Voice Separation and Enhancement”, In Proceedings of INTERSPEECH, Aug. 30, 2021, pp. 1124-1128. [cited by applicant]
Gulati, et al., “Conformer: Convolution-augmented Transformer for Speech Recognition”, In Repository of arXiv:2005.08100v1, May 16, 2020, 5 Pages. [cited by applicant]
Halimeh, et al., “Efficient Multichannel Nonlinear Acoustic Echo Cancellation Based on a Cooperative Strategy”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, May 4, 2020, pp… [cited by applicant]
Hao, et al., “SNR-Based Teachers-Student Technique for Speech Enhancement”, In Repository of arXiv:2005.14441v2, Oct. 29, 2020, 6 Pages. [cited by applicant]
Aggarwal, et al., “1329-1999—IEEE Standard Method for Measuring Transmission Performance of Handsfree Telephone Sets”, In Proceedings of IEEE, May 2, 2000, 94 Pages. [cited by applicant]
Hinton, et al., “Distilling the Knowledge in a Neural Network”, In Repository of arXiv:1503.02531v1, Mar. 9, 2015, 9 Pages. [cited by applicant]
Hochreiter, et al., “Long Short-Term Memory”, In Journal of Neural Computation, vol. 9, Issue 8, Nov. 15, 1997, pp. 1735-1780. [cited by applicant]
Hu, et al., “DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement”, In Repository of arXiv:2008.00264v1, Aug. 1, 2020, 5 Pages. [cited by applicant]
Ianniello, John P., “Time Delay Estimation Via Cross-Correlation in the Presence of Large Estimation Errors”, In Journal of IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 30, Issue 6, Dec. 6, 1982, … [cited by applicant]
Indenbom, et al., “Deep Model with Built-in Self-Attention Alignment for Acoustic Echo Cancellation”, In Repository of arXiv:2208.11308v1, Aug. 24, 2022, 5 Pages. [cited by applicant]
Isik, et al., “Single-Channel Multi-Speaker Separation using Deep Clustering”, In Proceedings of INTERSPEECH, Sep. 8, 2016, pp. 545-549. [cited by applicant]
What is Project Acoustics?, Retrieved From: https://web.archive.org/web/20220303165340/https://docs.microsoft.com/en-us/gaming/acoustics/what-is-acoustics, Apr. 26, 2021, 6 Pages. [cited by applicant]
Kim, et al., “Attention Wave-U-Net for Acoustic Echo Cancellation”, In Proceedings of INTERSPEECH, Oct. 25, 2020, pp. 3969-3973. [cited by applicant]
Kim, et al., “Test-Time Adaptation Toward Personalized Speech Enhancement: Zero-Shot Learning with Knowledge Distillation”, In Proceedings of IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, Oc… [cited by applicant]
Kingma, et al., “ADAM: A Method for Stochastic Optimization”, In Proceedings of 3rd International Conference on Learning Representations, May 7, 2015, 15 Pages. [cited by applicant]
Kobayashi, et al., “Implementation of Low-Latency Electrolaryngeal Speech Enhancement Based on Multi-Task CLDNN”, In Proceedings of 28th European Signal Processing Conference, Jan. 18, 2021, pp. 396-400. [cited by applicant]
Lee, et al., “DNN-Based Residual Echo Suppression”, In Proceedings of the Sixteenth Annual Conference of the International Speech Communication Association, Sep. 6, 2015, pp. 1775-1779. [cited by applicant]
Lin, et al., “A Low-Complexity Adaptive Echo Canceller for xDSL Applications”, In Journal of IEEE Transactions on Signal Processing, vol. 52, Issue 5, May 5, 2004, pp. 1461-1465. [cited by applicant]
Luo, et al., “Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation”, In Journal of IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 27, Issue 8, May 6, 2019, pp. 1256… [cited by applicant]
Ma, et al., “Acoustic Echo Cancellation by Combining Adaptive Digital Filter and Recurrent Neural Network”, In Repository of arXiv:2005.09237v1, May 19, 2020, 5 Pages. [cited by applicant]
Ma, et al., “EchoFilter: End-to-End Neural Network for Acoustic Echo Cancellation”, In Repository of arXiv:2105.14666v1, May 31, 2021, 5 Pages. [cited by applicant]
Naderi, et al., “An Open Source Implementation of ITU-T Recommendation P.808 with Validation”, In Proceedings of INTERSPEECH, Oct. 25, 2020, pp. 2862-2866. [cited by applicant]
Naderi, et al., “Subjective Evaluation of Noise Suppression Algorithms in Crowdsourcing”, In Proceedings of INTERSPEECH, Aug. 30, 2021, pp. 2132-2136. [cited by applicant]
Pandey, et al., “Dense CNN With Self-Attention for Time-Domain Speech Enhancement”, In Proceedings of IEEE/ACM Transactions on Audio, Speech, and Language Processing, Mar. 8, 2021, pp. 1270-1279. [cited by applicant]
De, et al., “Impact of Digital Surge During Covid-19 Pandemic: A Viewpoint on Research and Practice”, In Journal of Information Management, vol. 55, Jun. 9, 2020, 5 Pages. [cited by applicant]
Peng, et al., “ICASSP 2021 Acoustic Echo Cancellation Challenge: Integrated Adaptive Echo Cancellation with Time Alignment and Deep Learning-Based Residual Echo Plus Noise Suppression”, In Proceedings of IEEE Internatio… [cited by applicant]
Peng, et al., “Shrinking Bigfoot: Reducing Wav2vec 2.0 Footprint”, In Repository of arXiv:2103.15760v2, Apr. 1, 2021, 5 Pages. [cited by applicant]
Purin, et al., “AECMOS: A speech Quality Assessment Metric for Echo Impairment”, In Repository of arXiv:2110.03010v2, Oct. 8, 2021, 5 Pages. [cited by applicant]
Reddy, et al., “DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors”, In Repository of arXiv:2110.01763v4, Feb. 4, 2022, 5 Pages. [cited by applicant]
Reddy, et al., “INTERSPEECH 2021 Deep Noise Suppression Challenge”, In Repository of arXiv:2101.01902v3, Apr. 5, 2021, 5 Pages. [cited by applicant]
Reddy, et al., “The INTERSPEECH Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results”, In Proceedings of INTERSPEECH, Oct. 25, 2020, pp. 2492-2496. [cited by applicant]
Sanh, et al., “DistilBERT, A Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter”, In Repository of arXiv:1910.01108v1, Oct. 2, 2019, 5 Pages. [cited by applicant]
Sato, et al., “Should We Always Separate ?: Switching Between Enhanced and Observed Signals for Overlapping Speech Recognition”, In Proceedings of INTERSPEECH, Aug. 30, 2021, pp. 1149-1153. [cited by applicant]
Seo, et al., “Shortcut Connections based Deep Speaker Embeddings for End-to-End Speaker Verification System”, In Proceedings of INTERSPEECH System, vol. 13, No. 15, Sep. 15, 2019, pp. 2928-2932. [cited by applicant]
Sivaraman, et al., “Personalized Speech Enhancement through Self-Supervised Data Augmentation and Purification”, In Proceedings of INTERSPEECH, Aug. 30, 2021, pp. 2676-2680. [cited by applicant]
Sridhar, et al., “ICASSP 2021 Acoustic Echo Cancellation Challenge: Datasets, Testing Framework, and Results”, In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Jun. 6, 2021… [cited by applicant]
Sun, et al., “Explore Relative and Context Information with Transformer for Joint Acoustic Echo Cancellation and Speech Enhancement”, In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal P… [cited by applicant]
Taal, et al., “An Algorithm for Intelligibility Prediction of Time-Frequency Weighted Noisy Speech”, In Proceedings of IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, No. 7, Sep. 2011, pp. 2125-213… [cited by applicant]
Taherian, et al., “One Model to Enhance Them All: Array Geometry Agnostic Multi-Channel Personalized Speech Enhancement”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. … [cited by applicant]
Thakker, et al., “Fast Real-time Personalized Speech Enhancement: End-to-End Enhancement Network (E3Net) and Knowledge Distillation”, In Repository of arXiv:2204.00771v1, Apr. 2, 2022, 5 Pages. [cited by applicant]
Thiemann, et al., “The Diverse Environments Multi-channel Acoustic Noise Database (DEMAND): A Database of Multichannel Environmental Noise Recordings”, In Journal of of Meetings on Acoustics ICA2013, vol. 19, Issue 1, J… [cited by applicant]
Turc, et al., “Well-Read Students Learn Better: On the Importance of Pre-training Compact Models”, In Repository of arXiv:1908.08962v2, Sep. 25, 2019, 13 Pages. [cited by applicant]
Vaswani, et al., “Attention is All you Need”, In Proceedings of Advances in Neural Information Processing Systems, 2017, 11 Pages. [cited by applicant]
Wang, et al., “Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks”, In Proceedings of IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, Issue 6,… [cited by applicant]
Wang, et al., “Transformer-Based Acoustic Modeling for Hybrid Speech Recognition”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 9, 2020, pp. 6874-6878. [cited by applicant]
Wang, et al., “VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking”, In Proceedings of INTERSPEECH, Sep. 15, 2019, pp. 2728-2732. [cited by applicant]
Wang, et al., “VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition”, In Proceedings of INTERSPEECH, Oct. 25, 2020, pp. 2677-2681. [cited by applicant]
Watcharasupat, et al., “End-to-End Complex-Valued Multidilated Convolutional Neural Network for Joint Acoustic Echo Cancellation and Noise Suppression”, In Repository of arXiv:2110.00745v2, Oct. 11, 2021, 5 Pages. [cited by applicant]
Westhausen, et al., “Acoustic Echo Cancellation with the Dual-Signal Transformation LSTM Network”, In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Jun. 6, 2021, pp. 7138-7… [cited by applicant]
Williamson, et al., “Complex Ratio Masking for Monaural Speech Separation”, In Proceedings of IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 24, No. 3, Mar. 2016, pp. 483-492. [cited by applicant]
Wisdom, et al., “Differentiable Consistency Constraints for Improved Deep Speech Enhancement”, In Proceedings of the ICASSP IEEE International Conference on Acoustics, Speech and Signal Processing, May 12, 2019, pp. 900… [cited by applicant]
Xia, et al., “Weighted Speech Distortion Losses for Neural-Network-Based Real-Time Speech Enhancement”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, May 4, 2020, pp. 871-87… [cited by applicant]
Xiao, et al., “Single-channel Speech Extraction Using Speaker Inventory and Attention Network”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 17, 2019, pp. 86-90. [cited by applicant]
Xu, et al., “Can Model Compression Improve NLP Fairness?”, In Repository of arXiv:2201.08542v1, Jan. 21, 2022, 9 Pages. [cited by applicant]
Yin, et al., “PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement Network”, In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, No. 05, Apr. 3, 2020, pp. 9458-9465. [cited by applicant]
Yu, et al., “NeuralEcho: A Self-Attentive Recurrent Neural Network For Unified Acoustic Echo Suppression and Speech Enhancement”, In Repository of arXiv:2205.10401v1, May 20, 2022, 5 Pages. [cited by applicant]
Zhang, et al., “A Deep Learning Approach to Multi-Channel and Multi-Microphone Acoustic Echo Cancellation”, In Proceedings of INTERSPEECH, Aug. 30, 2021, pp. 1139-1143. [cited by applicant]
Zhang, et al., “A Robust and Cascaded Acoustic Echo Cancellation Based on Deep Learning”, In Proceedings of INTERSPEECH, Oct. 25, 2020, pp. 3940-3944. [cited by applicant]
Zhang, et al., “Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, Nov. 2, 2019, pp. 3713-37… [cited by applicant]
Zhang, et al., “Deep Learning for Joint Acoustic Echo and Noise Cancellation with Nonlinear Distortions”, In Proceedings of INTERSPEECH, Sep. 15, 2019, pp. 4255-4259. [cited by applicant]
Zhang, et al., “Multi-Scale Temporal Frequency Convolutional Network With Axial Attention for Speech Enhancement”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, May 23, 2022… [cited by applicant]
Zhang, et al., “Multi-Task Deep Residual Echo Suppression with Echo-Aware Loss”, In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, May 23, 2022, pp. 9127-9131. [cited by applicant]
Zhao, et al., “A Deep Hierarchical Fusion Network for Fullband Acoustic Echo Cancellation”, In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, May 23, 2022, pp. 9112-9116. [cited by applicant]
Zhou, et al., “ResNext and Res2Net Structures for Speaker Verification”, In Proceedings of IEEE Spoken Language Technology Workshop, Mar. 25, 2021, pp. 301-307. [cited by applicant]
“ITU-T P.808 : Subjective Evaluation of Speech Quality with a Crowdsourcing Approach”, Retrieved from: https://www.itu.int/itu-t/recommendations/rec.aspx?rec=13625&lang=en, Jun. 13, 2018, 28 Pages. [cited by applicant]