IP Library › Granted Patent US 12,272,371
Granted Patent B1
US 12,272,371 · App. 17/364,805 · Granted Apr 8, 2025

Real-time target speaker audio enhancement

Inventors: Ritwik Giri (Sunnyvale, CA); Shrikant Venkataramani (Champaign, IL); Jean-Marc Valin (Montreal, CA); Mehmet Umut Isik (Menlo Park, CA); Arvindh Krishnaswamy (Palo Alo, CA)
Assignee: Amazon Technologies, Inc.
G10L21/0364G06N20/00G10L21/013G10L21/038
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,272,371
App. No.
17/364,805
Granted
Apr 8, 2025
Kind
B1
Abstract

Real-time audio enhancement for a target speaker may be performed. An embedding of a sample of speaker audio is created using a trained neural network that performs voice identification. The embedding is then concatenated with the input features of a trained machine learning model for audio enhancement. The audio enhancement model can recognize and enhance a target speaker's speech in a real-time implementation, as the embedding is in the same feature space of the audio enhancement model.

Claims (44)

1. A system, comprising:

at least one processor; and

a memory, storing program instructions that when executed by the at least one processor, cause the at least one processor to implement an audio enhancement system, the audio enhancement system configured to:

receive audio data comprising speaker audio data via an interface for the audio enhancement system;

determine a plurality of input features for the audio data based on a representation of the audio data in an equivalent rectangular bandwidth scale;

obtain an embedding for a speaker generated from input features of an audio data sample for the speaker determined based on a representation of the audio data sample in the equivalent rectangular bandwidth scale;

apply a machine learning model trained to provide one or more modifications to enhance the speaker audio data within the audio data, wherein the machine learning model concatenates the embedding for the speaker with input features of the audio data to determine respective gain values for respective bands of the representation of the audio data in the equivalent rectangular bandwidth scale;

apply the one or more modifications to the audio data to generate an enhanced version of the audio data according to the respective gain values for the respective bands of the representation of the audio data in the equivalent rectangular bandwidth scale; and

send, via the interface of the audio enhancement system, the enhanced version of the audio data to a destination.

2. The system of claim 1 , wherein to obtain the embedding for the speaker, the audio enhancement system is configured to:

receive the audio data sample for the speaker via the interface; and

apply a speaker embedder network to the input features of the audio data sample to generate the embedding from a normalized last frame of the audio data sample output from the speaker embedder network.

3. The system of claim 1 , wherein the system further comprises an audio sensor that captures the audio data and wherein the destination is an audio-transmission service implemented as part of a provider network that transmits the enhanced version of the audio data to an audio playback device over a network connection.

4. The system of claim 1 , wherein the audio enhancement system is implemented as part of an audio-transmission service offered by a provider network, wherein the interface for the audio enhancement system supports receiving the audio data via a network connection, and wherein the destination is an audio playback device identified by the audio-transmission service for the audio data.

5. A method, comprising:

receiving audio data comprising speaker audio data via an interface for an audio enhancement system;

applying, by the audio enhancement system, a machine learning model trained to provide one or more modifications to enhance the speaker audio data within the audio data, wherein the machine learning model:

concatenates an embedding generated for a speaker with input features of the audio data to determine respective gain values for respective bands of a representation of the audio data in an equivalent rectangular bandwidth scale; and

wherein the input features of an audio data sample for the speaker used to generate the embedding are determined based on a further representation of the audio data sample in the same equivalent rectangular bandwidth scale; and

providing, by the audio enhancement system, an enhanced version of the audio data generated, based on the one or more modifications to enhance the speaker audio data according to the respective gain values for the respective bands of the representation of the audio data in the equivalent rectangular bandwidth scale.

6. The method of claim 5 , further comprising:

receiving the audio data sample for the speaker via the interface; and

applying a speaker embedder network to the input features of the audio data sample to generate the embedding from a normalized last frame of the audio data sample output from the speaker embedder network.

7. The method of claim 6 , further comprising prompting the speaker to provide the audio data sample via the interface.

8. The method of claim 7 , further comprising rejecting a previously provided audio data sample for the speaker before prompting the speaker to provide the audio data sample.

9. The method of claim 5 , wherein applying the machine learning model trained to provide one or more modifications to enhance the speaker audio data within the audio data further comprises determining an activity estimate for the speaker for respective frames of the audio data.

10. The method of claim 5 , further comprising receiving, via the interface for the audio enhancement system, the embedding for the speaker.

11. The method of claim 5 , wherein the audio data is captured along with corresponding video data that is provided to a same destination as the enhanced version of the audio data.

12. The method of claim 5 , wherein providing the enhanced version of the audio data comprises storing the enhanced version of the audio data to a data storage service offered by a provider network.

13. The method of claim 5 , wherein the audio enhancement system is implemented as part of a device that includes an audio sensor that captured the audio data, and wherein providing the enhanced version of the audio data comprises sending the enhanced version of the audio data to an audio-transmission service implemented as part of a provider network that transmits the enhanced version of the audio data to an audio playback device over a network connection.

14. One or more non-transitory, computer-readable storage media, storing program instructions that when executed on or across one or more computing devices cause the one or more computing devices to implement:

receiving audio data comprising speaker audio data via an interface for an audio enhancement system;

causing application of a machine learning model trained to provide one or more modifications to enhance the speaker audio data within the audio data, wherein the machine learning model:

concatenates an embedding generated for a speaker with input features of the audio data to determine respective gain values for respective bands of a representation of the audio data in an equivalent rectangular bandwidth scale; and

wherein the input features of an audio data sample for the speaker used to generate the embedding are determined based on a further representation of the audio data sample in the same equivalent rectangular bandwidth scale; and

sending, by the audio enhancement system, an enhanced version of the audio data generated, based on the one or more modifications to enhance the speaker audio data according to the respective gain values for the respective bands of the representation of the audio data in the equivalent rectangular bandwidth scale, to a destination.

15. The one or more non-transitory, computer-readable storage media of claim 14 , storing further instructions that when executed on or across the one or more computing devices cause the one or more computing devices to further implement:

receiving the audio data sample for the speaker via the interface; and

applying a speaker embedder network to the input features of the audio data sample to generate the embedding from a normalized last frame of the audio data sample output from the speaker embedder network.

16. The one or more non-transitory, computer-readable storage media of claim 15 , storing further instructions that when executed on or across the one or more computing devices cause the one or more computing devices to further implement prompting the speaker to provide the audio data sample via the interface.

17. The one or more non-transitory, computer-readable storage media of claim 16 , storing further instructions that when executed on or across the one or more computing devices cause the one or more computing devices to further implement rejecting a previously provided audio data sample for the speaker before prompting the speaker to provide the audio data sample.

18. The one or more non-transitory, computer-readable storage media of claim 14 , wherein, in applying the machine learning model trained to provide one or more modifications to enhance the speaker audio data within the audio data, the program instructions cause the one or more computing devices to determining an activity estimate for the speaker for respective frames of the audio data.

19. The one or more non-transitory, computer-readable storage media of claim 14 , wherein the audio enhancement system is implemented as part of a device that includes an audio sensor that captured the audio data, and wherein providing the enhanced version of the audio data comprises sending the enhanced version of the audio data to an audio-transmission service implemented as part of a provider network that transmits the enhanced version of the audio data to an audio playback device over a network connection.

20. The one or more non-transitory, computer-readable storage media of claim 14 , wherein the audio data is captured along with corresponding video data that is provided to a same destination as the enhanced version of the audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2021
From: GIRI, RITWIK; VENKATARAMANI, SHRIKANT; VALIN, JEAN-MARC; ISIK, MEHMET UMUT; KRISHNASWAMY, ARVINDH
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 058150/0936 →
References Cited (64)
US 7027591B2 · Cairns · 2006 [cited by applicant]
US 8416946B2 · Chhetri et al. · 2013 [cited by applicant]
US 10904396B2 · Hera et al. · 2021 [cited by applicant]
US 11521637B1 · Valin et al. · 2022 [cited by applicant]
US 20180115824A1 · Cassidy · 2018 [cited by applicant]
US 20190201657A1 · Popelka · 2019 [cited by examiner]
US 20190318755A1 · Tashev · 2019 [cited by applicant]
US 20200066296A1 · Sargsyan · 2020 [cited by applicant]
US 20200152179A1 · van Hout · 2020 [cited by examiner]
US 20210125625A1 · Huang · 2021 [cited by applicant]
US 20220122597A1 · Ji · 2022 [cited by examiner]
US 20220335953A1 · Rikhye · 2022 [cited by examiner]
US 20230419984A1 · Uhle · 2023 [cited by examiner]
Ding Liu, et al., “Experiments on Deep Learning for Speech Denoising”, In Proceedings of Fifteenth Annual Conference of the International Speech Communication Association, 2014, pp. 1-5. [cited by applicant]
Yong Xu, et al., “A Regression Approach to Speech Enhancement Based on Deep Neural Networks”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 23, No. 1, Jan. 2015, pp. 7-19. [cited by applicant]
Ke Tan, et al., “A Convolutional Recurrent Neural Network for Real-Time Speech Enchancement”, Interspeech 2018, Sep. 2-6, 2018, pp. 1-5. [cited by applicant]
Arun Narayanan, et al., “Ideal Ratio Mask Estimation Using Deep Neural Networks for Robust Speech Recognition”, IEEE, In Proceedings of International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2013,… [cited by applicant]
Yan Zhao, et al., “DNN-Based Enhancement of Noisy and Reverberant Speech”, IEEE, In Proceedings of International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016, pp. 6525-6529. [cited by applicant]
Donald S. Williamson, et al., “Complex Ratio Masking for Monaural Speech Separation”, IEEE/ACM Transactions on Audio Speech Language Processing, 24(3), Mar. 2016, pp. 483-492. [cited by applicant]
Santiago Pascual, et al., “SEGAN: Speech Enhancement Generative Adversarial Network”, arXiv:1703.09452v3, Jun. 9, 2017, pp. 1-5. [cited by applicant]
Dario Rethage, et al., “A Wavenet for Speech Denoising”, arXiv:1706.07162v3, Jan. 31, 2018, pp. 1-11. [cited by applicant]
Craig Macartney, et al., “Improved Speech Enhancement with the Wave-U-Net”, arXiv:1811.11307v1, Nov. 27, 2018, pp. 1-5. [cited by applicant]
Jean-Marc Valin, “A Hybrid DSP/Deep Learning Approach to Real-Time Full-Band Speech Enhancement”, arXiv:1709.08243v3, May 31, 2018, pp. 1-5. [cited by applicant]
John P. Princen, et al., “Analysis/Synthesis Filter Bank Design Based on Time Domain Aliasing Cancellation”, IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. ASSP-34, No. 5, Oct. 1986, pp. 1153-1161. [cited by applicant]
Hedwig Gockel, et al., “Asymmetry of masking between complex tones and noise: Partial loudness”, The Journal of the Acoustical Society of America, 114(1), Jul. 2003, pp. 349-360. [cited by applicant]
D. Talkin. A robust algorithm for pitch tracking (RAPT). In Speech Coding and Synthesis, chapter 14, Elsevier Science, 1995, pp. 495; 497-518. [cited by applicant]
Ted Painter, et al., “Perceptual Coding of Digital Audio”, in Proceedings of the IEEE, vol. 88, No. 4, Apr. 2000, pp. 451-513. [cited by applicant]
Kyunghyun Cho, et al., “On the Properties of Neural Machine Translation: Encoder-Decorder Approaches”, arXiv:1409.1259v2, Oct. 7, 2014, pp. 1-9. [cited by applicant]
Yan Zhao, et al., “Late Reverberation Suppression Using Recurrent Neural Networks With Long Short-Term Memory”, In Proceedings of International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2018,… [cited by applicant]
Hakan Erdogan, et al., “Investiagtions on Data Augmentation and Loss Functions for Deep Learning Based Speech-Background Separation”, Interspeech 2018, Sep. 2-6, 2018, pp. 1-5. [cited by applicant]
Juin-Hwey Chen, et al., “Adaptive Postfiltering for Quality Enhancement of Coded Speech”, in IEEE Transactions on Speech and Audio Processing, vol. 3, No. 1, Jan. 1995, pp. 59-71. [cited by applicant]
Valentini Botinhao, et al., “Investigating RNN-based speech enhancement methods for noise-robust Text-to-Speech”, In Proceedings of ISCA Speech Synthesis Workshop (SSW), 2016, pp. 146-152. [cited by applicant]
U.S. Appl. No. 17/668,297, filed Feb. 9, 2022, Jean-Marc Valin, et al. [cited by applicant]
Yangyang Xia, et al., Weighted Speech Distortion Losses for Neural-Network-Based Real-Time Speech Enhancement, arXiv:2001.10601v2. Feb. 12, 2020, pp. 1-5. [cited by applicant]
U.S. Appl. No. 17/037,498, filed Sep. 29, 2020, Jean-Marc Valin et al. [cited by applicant]
U.S. Appl. No. 17/037,515, filed Sep. 29, 2020, Mehmet Umut Isik et al. [cited by applicant]
Jean-Marc Valin, et al., “A PErceptually-Motivated Approach for Low-Complexity, Rea-Time Enhancement of Fullband Speech”, arXiv:2008.04259v2, Aug. 27, 2020, pp. 1-5. [cited by applicant]
Yi Luo, “TASNET: Time-Domain Audio Separation Network for Real-Time, Single-Channel Speech Separation”, arXiv:1711.00541v2, Apr. 8, 2018, pp. 1-5. [cited by applicant]
John R. Hershey, et al., “Deep clustering: Discriminative embeddings for segmentation and separation”, arXiv:1508.04306v1, Aug. 18, 2015, pp. 1-10. [cited by applicant]
Morten Holbaek, et al., “Multi-talker Speech Separation with Utterance-level Permutation Invariant Trainingof Deep Recurrent Neural Networks”, arXiv:1703.06284v2, Jul. 11, 2017, pp. 1-12. [cited by applicant]
Yi Luo, et al., “Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation”, IEEE/ACM Transactions on Audio, Speech and Language Processing, vol. 27, No. 8, Aug. 2019, pp. 1256-1266. [cited by applicant]
Efthymios Tzinis, et al., “Unspervised Deep Clustering for Source Separation: Direct Learning From Mixtures Using Spatial Informaiton”, arXiv:1811.01531v2, Nov. 9, 2018, pp. 1-5. [cited by applicant]
Thilo von Neumann, et al., “All-Neural Online Source Separation, Counting, and Diarization for Meeting Analysis”, arXiv:1902.07881v1, Feb. 21, 2019, pp. 1-5. [cited by applicant]
Desh Raj, et al., “Integration of Speech Separation, Diarization, and Recognition for Multi-Speaker Meetings: System Description, Comparison, and Anaylsis”, arXiv:2011.02014v1, Nov. 3, 2020, pp. 1-8. [cited by applicant]
Quan Wang, et al., “VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking”, arXiv:1810.04826v6, Jun. 19, 2019, pp. 1-5. [cited by applicant]
Lin Wan, et al., “Generalized End-to-End Loss for Speaker Verification”, arXiv:1710.10467v5, Nov. 9, 2020, pp. 1-5. [cited by applicant]
Seongkyu Mun, et al. “The Sound of My Voice: Speaker Representation Loss for Target Voice Separation”, arXiv:1911.02411v2, Feb. 27, 2020, pp. 1-5. [cited by applicant]
Rongzhi Gu, et al., “Neural Spatial Filter: Target Speaker Speech Separation Assisted with Directional Information”, Interspeech 2019, Sep. 15-19, 2019, Graz, Austria, pp. 1-5. [cited by applicant]
Tingle Li, et al., “Atss-Net: Target Speaker Separation via Attention-based Neural Network”, arXiv:2005.09200v1, May 19, 2020, pp. 1-5. [cited by applicant]
Xiong Xiao, et al., “Speech Separation Using Speaker Inventory”, IEEE, ASRU 2019, pp. 230-236. [cited by applicant]
Quan Wang, et al., “VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition”, arXiv:2009-04323v1, Sep. 9, 2020, pp. 1-5. [cited by applicant]
Chandan K. A. Reddy, et al., “The Interspeech 2020 Deep Noise Suppression Challenge: Datasets, Subjective Speech Quality and Testing Framework”, arXiv preprint arXiv:2001.08662, 2020, pp. 1-5. [cited by applicant]
Aonan Zhang, et al., “Fully Surpervised Speaker Diarization”, arXiv:1810.04719v7, Feb. 19, 2019, pp. 1-5. [cited by applicant]
Quan Wang, et al., “Speaker Diarization With LSTM”, arXiv:1710.10468v6, Dec. 14, 2018, pp. 1-5. [cited by applicant]
Ye Jia, et al., “Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis”, arXiv:1806.04558v4, Jan. 2, 2019, pp. 1-15. [cited by applicant]
Kaizhi Qian, et al., “AutoVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss”, in Proceedings of the 36th International Conference on Machine Learning, PLMR 97, 2019, pp. 1-10. [cited by applicant]
Jean-Marc Valin, et al., “LPCNET: Improving Neural Speech Synthesis Through Linear Prediction”, arXiv:1810.11846v2, Feb. 19, 2019, pp. 1-5. [cited by applicant]
Sebastian Braun, et al., “Data augmentation and loss normalization for deep noise suppression”, arXiv:2008.06412v2, Sep. 24, 2020, pp. 1-8. [cited by applicant]
Shaojin Ding, et al., “Personal VAD: Speaker-Conditioned Voice Activity Detection”, arXiv:1908.04284v4, Apr. 8, 2020, pp. 1-7. [cited by applicant]
Joon Son Chung, et al., “VoxCeleb2: Deep Speaker Recognition”, arXiv: 1806.05622v2, Jun. 27, 2018, pp. 1-6. [cited by applicant]
Arsha Nagrani, et al., “VoxCeleb: a large-scale speaker identification dataset”, arXiv:1706.08612v2, May 30, 2018, pp. 1-6. [cited by applicant]
Vassil Panayotov, et al., “Librispeech: an ASR Corpus Based on Public Domain Audio Books”, 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2015, pp. 1-5. [cited by applicant]
Umut Isik, et al., “PoCoNet: Better Speech Enhancement with Frequency-Positional Embeddings, Semi-Supervised Conversational Data, and Biased Loss”, arXiv:2008-04470v1, Aug. 11, 2020, pp. 1-5. [cited by applicant]
Joachim Thiemann, et al., “Demand: a collection of multi-channel recordings of acoustic noise in diverse environments”, Version 1.0, Jun. 9, 2013, In Proceedings Meetings Acoustic, pp. 1-6. [cited by applicant]
Cited By (2)
US 12,482,482 US 12,586,590