IP Library › Granted Patent US 12,499,903
Granted Patent B2
US 12,499,903 · App. 18/054,856 · Granted Dec 16, 2025

Neural network training for speech enhancement

Inventors: Saeed Mosayyebpour Kaskari (Irvine, CA); Atabak Pouya (Irvine, CA)
Assignee: Synaptics Incorporated
G10L25/30G06N3/082G10L21/0232G10L25/18G10L25/21
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,903
App. No.
18/054,856
Granted
Dec 16, 2025
Kind
B2
Abstract

A method of training neural networks may include receiving a sequence of audio frames, and mapping a first audio frame in the sequence of audio frames to a first output frame based on a neural network. The first output frame may represent a noise-invariant component of the first audio frame. The method may also include determining a first loss value based on differences between the first output frame and a first ground truth frame. The method may include mapping the first audio frame to a second output frame based on the neural network. The second output frame may represent a noise-variant component of the first audio frame. The method may further include determining a second loss value based on differences between the second output frame and a second ground truth frame, and updating the neural network based at least in part on the first and second loss values.

Claims (75)

1 . A method of training neural networks, comprising:

receiving a sequence of audio frames representing an audio signal;

estimating a speech component of a first audio frame in the sequence of audio frames using a first head of a multi-head neural network;

mapping, using the first head of the multi-head neural network, the first audio frame to a first output frame representing the estimated speech component of the first audio frame;

receiving a first ground truth frame including a known speech component of the first audio frame;

determining a first loss value based on differences between the first output frame and the first ground truth frame;

estimating a gain associated with a signal-to-noise ratio (SNR) of the first audio frame using a second head of the multi-head neural network;

mapping, using the second head of the multi-head neural network, the first audio frame to a second output frame representing the estimated gain of the first audio frame;

receiving a second ground truth frame including a known gain of the first audio frame;

determining a second loss value based on differences between the second output frame and the second ground truth frame;

determining a total loss value based on the first loss value and the second loss value; and

iteratively updating the multi-head neural network until the total loss value converges to a predetermined threshold amount for inferring speech from real-time audio signals.

2 . The method of claim 1 , wherein the mapping of the first audio frame to the second output frame is further based on a phase of the first audio frame.

3 . The method of claim 2 , wherein the mapping of the first audio frame to the second output frame is further based on a magnitude of the first audio frame.

4 . The method of claim 1 , wherein the second output frame further represents a ratio of an amount of energy in the speech component of the first audio frame relative to an amount of energy of the first audio frame.

5 . The method of claim 1 , wherein the received first audio frame is normalized based on a spectrum magnitude.

6 . The method of claim 1 , wherein the first audio frame includes a plurality of sub-frames, and wherein each of the plurality of sub-frames is associated with a respective frequency bin.

7 . The method of claim 1 , wherein determining the total loss value comprises:

summing at least the first loss value and the second loss value.

8 . The method of claim 7 , wherein the updating of the multi-head neural network further comprises:

determining a set of parameters for the multi-head neural network that minimizes the total loss value.

9 . The method of claim 1 , further comprising:

estimating a speech component of a second audio frame in the sequence of audio frames using the first head of a multi-head neural network;

mapping, using the first head of the multi-bead neural network, the second audio frame to a third output frame representing the estimated speech component of the second audio frame;

receiving a third frame including a known speech component of the second audio frame;

determining a third loss value based on differences between the third output frame and the third ground truth frame;

estimating a gain associated with an SNR of the second audio frame using the second head of the multi-head neural network;

mapping, using the second head of the multi-head neural network, the second audio frame to a fourth output frame representing the estimated gain of the second audio frame;

receiving a fourth ground truth frame including a known gain of the second audio frame; and

determining a fourth loss value based on differences between the fourth output frame and the fourth ground truth frame, wherein the multi-head neural network is further updated based on the third loss value and the fourth loss value.

10 . A machine learning system comprising:

a processing system; and

a memory storing instructions that, when executed by the processing system, cause the machine learning system to:

receive a sequence of audio frames representing an audio signal;

estimate a speech component of a first audio frame in the sequence of audio frames using a first head of a multi-head neural network;

map, using the first head of the multi-head neural network, the first audio frame to a first output frame representing the estimated speech component of the first audio frame;

receive a first ground truth frame including a known speech component of the first audio frame;

determine a first loss value based on differences between the first output frame and the first ground truth frame;

estimate a gain associated with a signal-to-noise ratio (SNR) of the first audio frame using a second head of the multi-head neural network;

map, using the second head of the multi-head neural network, the first audio frame to a second output frame representing the estimated gain of the first audio frame;

receive a second ground truth frame including a known gain of the first audio frame;

determine a second loss value based on differences between the second output frame and the second ground truth frame;

determine a total loss value based on the first loss value and the second loss value; and

iteratively update the multi-head neural network until the total loss value converges to a predetermined threshold amount for inferring speech from real-time audio signals.

11 . The machine learning system of claim 10 , wherein the mapping of the first audio frame to the second output frame is further based on a phase of the first audio frame.

12 . The machine learning system of claim 11 , wherein the mapping of the first audio frame to the second output frame is further based on a magnitude of the first audio frame.

13 . The machine learning system of claim 10 , wherein the second output frame further represents a ratio of an amount of energy in the speech component of the first audio frame relative to an amount of energy of the first audio frame.

14 . The machine learning system of claim 10 , wherein the received first audio frame is normalized based on a spectrum magnitude.

15 . The machine learning system of claim 10 , wherein the first audio frame includes a plurality of sub-frames, and wherein each of the plurality of sub-frames is associated with a respective frequency bin.

16 . The machine learning system of claim 10 , wherein determining the total loss value further causes the machine learning system to:

sum at least the first loss value and the second loss value.

17 . The machine learning system of claim 16 , wherein the updating of the multi-head neural network further causes the machine learning system to:

determine a set of parameters for the multi-head neural network that minimizes the total loss value.

18 . The machine learning system of claim 10 , wherein execution of the instructions further causes the machine learning system to:

estimate a speech component of a second audio frame in the sequence of audio frames using the first head of a multi-bead neural network;

map, using the first head of the multi-head neural network, the second audio frame to a third output frame representing the estimated speech component of the second audio frame;

receive a third ground truth frame including a known speech component of the second audio frame;

determine a third loss value based on differences between the third output frame and the third ground truth frame;

estimate a gain associated with an SNR of the second audio frame using the second head of the multi-head neural network;

map, using the second head of the multi-head neural network, the second audio frame to a fourth output frame representing the estimated gain of the second audio frame;

receive a fourth ground truth frame including a known gain of the second audio frame; and

determine a fourth loss value based on differences between the fourth output frame and the fourth ground truth frame, wherein the multi-head neural network is further updated based on the third loss value and the fourth loss value.

19 . A method of training neural networks, comprising:

receiving a sequence of audio frames representing an audio signal;

estimating a speech component of a first audio frame in the sequence of audio frames using a first head of a multi-head neural network;

mapping, using the first head of the multi-head neural network, the first audio frame to a first output frame representing the estimated speech component of the first audio frame;

receiving a first ground truth frame including a known speech component of the first audio frame;

determining a first loss value based on differences between the first output frame and the first ground truth frame; and

iteratively updating the first neural network until one or more convergence criteria associated with the first loss value are met.

20 . The method of claim 19 , further comprising:

estimating a gain associated with a signal-to-noise ratio (SNR) of the first audio frame using a second head of the multi-head neural network;

mapping, using the second head of the multi-head neural network, the first audio frame to a second output frame representing the estimated gain of the first audio frame;

receiving a second ground truth frame including a known gain of the first audio frame;

determining a second loss value based on differences between the second output frame and the second ground truth frame; and

iteratively updating the second neural network until one or more convergence criteria associated with the second loss value are met.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 11, 2022
From: MOSAYYEBPOUR KASKARI, SAEED; POUYA, ATABAK
To: SYNAPTICS INCORPORATED
Reel/Frame 061744/0351 →
Continuity (1)
Related Publication 20240170008A1 · May 23, 2024
References Cited (51)
US 9064498B2 · Uhle · 2015 [cited by examiner]
US 10529349B2 · Le Roux · 2020 [cited by examiner]
US 10672414B2 · Tashev · 2020 [cited by examiner]
US 11817111B2 · Fejgin · 2023 [cited by examiner]
US 20160111108A1 · Erdogan · 2016 [cited by examiner]
WO WO2024051676A1 · 2024 [cited by examiner]
Wang, DeLiang, and Jitong Chen. “Supervised speech separation based on deep learning: An overview.” IEEE/ACM transactions on audio, speech, and language processing 26.10 (2018): 1702-1726. (Year: 2018). [cited by examiner]
Han, Kun, et al. “Learning spectral mapping for speech dereverberation and denoising.” IEEE/ACM Transactions on Audio, Speech, and Language Processing 23.6 (2015): 982-992. (Year: 2015). [cited by examiner]
Afouras et al., “The Conversation: Deep Audio Visual Speech Enhancement,” Proc. Interspeech 2018, pp. 3244-3248, 2018. [cited by applicant]
Arjovsky et al., “Unitary Evolution Recurrent Neural Networks,” in International Conference on Machine Learning, pp. 1120-1128, 2016. [cited by applicant]
Choi et al., “Phase-Aware Speech Enhancement with Deep Complex U-Net,” International Conference on Learning Representations, pp. 1-20, 2019. [cited by applicant]
Cogswell et al., “Reducing Overfitting in Deep Networks by Decorrelating Representations,” arXiv preprint arXiv:1511.06068, pp. 1-11, 2015. [cited by applicant]
Ephrat et al., “Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation,” arXiv preprint arXiv:1804.03619 v2, pp. 1-11, 2018. [cited by applicant]
Erdogan et al., “Phase-Sensitive and Recognition-Boosted Speech Separation Using Deep Recurrent Neural Networks,” in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on, pp. 708-712, … [cited by applicant]
Germain et al., “Speech Denoising with Deep Feature Losses,” arXiv preprint arXiv: 1806.10522 v2, pp. 1-6, 2018. [cited by applicant]
Glorot et al., “Understanding the Difficulty of Training Deep Feedforward Neural Networks,” in Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, pp. 249-256, 2010. [cited by applicant]
Grais et al., “Single-Channel Audio Source Separation Using Deep Nerual Network Ensembles,” in Audio Engineering Society Convention 140, Audio Engineering Society, pp. 1-7, 2016. [cited by applicant]
Griffin et al., “Signal Estimation from Modified Short-Time Fourier Transform,” IEEE Transactions on Acoustics, Speech, and Signal Processing, 32(2):236-243, 1984. [cited by applicant]
Huang et al., “Deep Learning for Monaural Speech Separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on, pp. 1562-1566, IEEE, 2014. [cited by applicant]
Jansson et al., “Singing Voice Separation with Deep U-Net Convolutional Networks,” in ISMIR, pp. 745-751, 2017. [cited by applicant]
Kim et al., “NSML: Meet the Mlaas Platform with a Real-World Case Study,” arXiv preprint arXiv:1810.09957, pp. 1-11, 2018. [cited by applicant]
Lee et al., “Fully Complex Deep Neural Network for Phase-Incorporating Monaural Source Separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 281-285, IEEE, 2017a. [cited by applicant]
Lee et al., “Discriminative Training of Complex-Valued Deep Recurrent Neural Network for Singing Voice Separation,” In Proceedings of the 2017 ACM on Multimedia Conference, pp. 1327-1335, ACM, 2017b. [cited by applicant]
Maas et al., “Rectifier Nonlinearities Improve Neural Network Acoustic Models,” in Proc. ICML, vol. 30, pp. 1-6, 2013. [cited by applicant]
Nugraha et al., “Multichannel Audio Source Separation with Deep Neural Networks,” IEEE/ACM Trans. Audio, Speech & Language Processing, 24(9): 1652-1664, 2016. [cited by applicant]
Pascual et al., “Segan: Speech Enhancement Generative Adversarial Network,” In Proc. Interspeech, 2017, pp. 3642-3646, 2017. [cited by applicant]
Perraudin et al., “A Fast Griffin-Lim Algorithm,” in Applications of Signal Processing to Audio and Acoustics (WASPAA), 2013 IEEE Workshop on, pp. 1-4, IEEE, 2013. [cited by applicant]
Rethage et al., “A Wavenet for Speech Denoising,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1-11, 2018. [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation,” in International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 234-241, 2015. [cited by applicant]
Roux et al., “Phase-book and Friends: Leveraging Discrete Representations for Source Separation,” arXiv preprint arXiv:1810.01395v1, pp. 1-12, 2018. [cited by applicant]
Scalart et al., “Speech Enhancement Based on a Priori Signal to Noise Estimation,” in Acoustics, Speech, and Signal Processing, 1996, ICASSP-96, Conference Proceedings, 1996 IEEE International Conference on, vol. 2, pp.… [cited by applicant]
Soni et al., “Time-Frequency Masking-Based Speech Enhancement Using Generative Adversarial Network,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5039-5043, 2018. [cited by applicant]
Stoller et al., “Wave-u-net: A Multi-scale Neural Network for End-to-End Audio Source Separation,” in ISMIR, pp. 334-340, 2018. [cited by applicant]
Sung et al., “NSML: A Machine Learning Platform that Enables You to Focus on Your Models,” arXiv preprint arXiv:1712.05902v1, pp. 1-8, 2017. [cited by applicant]
Takahashi et al., “Phasenet: Discretized Phase Modeling with Deep Neural Networks for Audio Source Separation,” Proc. Interspeech 2018, pp. 2713-2717, 2018a. [cited by applicant]
Takahashi et al., Mmdenselstm: An efficient Combination of Convolutional and Recurrent Neural Networks for Audio Source Separation, arXiv preprint arXiv:1805.02410, 2018b. [cited by applicant]
Thiemann et al., “The Diverse Environments Multi-Channel Acoustic Noise Database: A Database of Multichannel Environmental Noise Recordings,” The Journal of the Acoustical Society of America, 133(5):3591, 2013. [cited by applicant]
Trabelsi et al., “Deep Complex Networks,” in International Conference on Learning Representations, pp. 1-19, 2018. [cited by applicant]
Veaux et al., “The Voice Bank Corpus: Design, Collection and Data Analysis of a Large Regional Accent Speech Database,” in Oriental COCOSDA Held Jointly with 2013 Conference on Asian Spoken Language Research and Evaluat… [cited by applicant]
Venkataramani et al., “Adaptive Front-Ends for End-to-End Source Separation,” in Workshop Machine Learning for Audio Signal Processing at NIPS (ML4Audio@NIPS17), pp. 1-5, 2017. [cited by applicant]
Vincent et al., “Performance Measurement in Blind Audio Source Separation,” IEEE Transactions on Audio, Speech, and Language Processing, 14(4):1462-1469, 2006. ISSN 1558-7916.doi:10.1109/TSA.2005.858005. [cited by applicant]
Wang, “Deep Learning Reinvents the Hearing Aid,” IEEE Spectrum, 54(3): 32-37, 2017. [cited by applicant]
Wang et al., “On Training Targets for Supervised Speech Separation,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), 22 (12):1849-1858, 2014. [cited by applicant]
Wang et al., “Multi-Channel Deep Clustering: Discriminative Spectral and Spatial Embeddings for Speaker-Independent Speech Separation,” Mitsubishi Electric Research Laboratories, pp. 1-6, 2018. [cited by applicant]
Wang et al., “Oracle Performance Investigation of the Ideal Masks,” in Acoustic Signal Enhancement (IWAENC), 2016 EEE International Workshop on, pp. 1-5, IEEE, 2016. [cited by applicant]
Weninger et al., “Speech Enhancement with LSTM Recurrent Nerual Networks and Its Application to Noise-Robust ASR,” in International Conference on Latent Variable Analysis and Signal Separation, pp. 91-99, Springer, 2015. [cited by applicant]
Williamson et al., “Complex Ratio Masking for Monaural Speech Separation,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), 24(3):483-492, 2016. [cited by applicant]
Wisdom et al., “Full-Capacity Unitary Recurrent Neural Networks,” in Advances in Neural Information Processing Systems, pp. 4880-4888, 2016. [cited by applicant]
Xu et al., A Regression Approach to Speech Enhancement Based on Deep Neural Networks, IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), 23(1):7-19, 2015. [cited by applicant]
Yegnanarayana et al., “Significance of Group Delay Functions in Spectrum Estimation,” IEEE Transactions on Signal Processing, 40(9):2281-2289, 1992. [cited by applicant]
Yu et al., “Permutation Invariant Training of Deep Models for Speaker-Independent Multi-talker Speech Separation,” in Acoustics, Speech and Signal Processing (ICASSP), arXiv:1607.00325v2, pp. 241-245, IEEE, 2017. [cited by applicant]