IP Library Granted Patent US 12,266,346
Granted Patent B2
US 12,266,346 · App. 17/390,788 · Granted Apr 1, 2025

Noisy far-field speech recognition

Inventors: Dading Chong (Hangzhou, CN); Zhaoyi Liu (Hangzhou, CN); Vijay Parthasarathy (San Jose, CA); Xiao Song (Hangzhou, CN)
Assignee: Zoom Communications, Inc.
G10L15/063G10L15/02G10L25/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,266,346
App. No.
17/390,788
Granted
Apr 1, 2025
Kind
B2
Abstract

The accuracy of automatic speech recognition (ASR) tasks is improved using trained models. A speech recognition model is applied in a noisy environment where speech is spoken at a distance from the microphones. The techniques may include extracting speech features, data augmentation by adding feature perturbation, and/or a multi-domain end-to-end speech recognition model. In some implementations, the described technology includes using a teacher-group knowledge distillation strategy to train a deep end-to-end speech recognition model on original speech samples and the sample speech augmentation of the original speech samples, that outputs recognized text transcriptions corresponding to speech detected in the original speech samples and the sample speech augmentation.

Claims (56)

1. A method comprising:

training a model using audio recordings from noise scenarios in a set of training data;

decomposing a training signal from the set of training data into a message component and a noise component;

scaling the noise component by a random scale factor to obtain a scaled noise, wherein the random scale factor is a power with a base that is a constant and an exponent that includes a random variable;

adding the scaled noise to the message component to obtain a perturbed audio signal that is included in the set of training data;

training a first teacher model using a first subset of the set of training data associated with a first noise scenario of the noise scenarios;

training a second teacher model using a second subset of the set of training data associated with a second noise scenario of the noise scenarios; and

training a student model using soft labels output from the first teacher model and soft labels output from the second teacher model.

2. The method of claim 1 , wherein training the student model using soft labels output from the second teacher model comprises:

determining a label for a training signal from the set of training data as a linear interpolation of a soft label from the second teacher model and a hard label for the training signal.

3. The method of claim 1 , wherein training the student model using soft labels output from the second teacher model comprises:

randomly selecting a training signal from the set of training data;

identifying a noise scenario associated with the selected training signal; and

determining a label for the selected training signal as a linear interpolation of a hard label for the training signal and a soft label from a teacher model trained using a subset of the training data associated with the identified noise scenario.

4. The method of claim 1 , wherein the first subset of the training data associated with the first noise scenario is based on audio recordings from streets, and the second subset of the training data associated with the second noise scenario is based on audio recordings from rooms inside buildings.

5. The method of claim 1 ,

wherein the random scale factor is chosen from a range uniformly sampled in [−8 dB, −1 dB).

6. The method of claim 1 , wherein the message component is an audio signal recorded with a microphone near a desired audio source while the training signal is recorded with a microphone far from the desired audio source.

7. The method of claim 1 , wherein decomposing the training signal from the set of training data into the message component and the noise component comprises:

applying feature extraction, including a log-mel filter bank, to the training signal and to the message component; and

subtracting features of the message component from features of the training signal to obtain features of the noise component.

8. A system comprising:

a network interface,

a processor, and

a memory, wherein the memory stores instructions executable by the processor to:

train a model using audio recordings from noise scenarios in a set of training data;

decompose a training signal from the set of training data into a message component and a noise component;

scale the noise component by a random scale factor to obtain a scaled noise, wherein the random scale factor is a power with a base that is a constant and an exponent that includes a random variable;

add the scaled noise to the message component to obtain a perturbed audio signal that is included in the set of training data;

train a first teacher model using a first subset of the set of training data associated with a first noise scenario of the noise scenarios;

train a second teacher model using a second subset of the set of training data associated with a second noise scenario of the noise scenarios; and

train a student model using soft labels output from the first teacher model and soft labels output from the second teacher model.

9. The system of claim 8 , wherein the memory stores instructions executable by the processor to:

determine a label for a training signal from the set of training data as a linear interpolation of a soft label from the second teacher model and a hard label for the training signal.

10. The system of claim 8 , wherein the memory stores instructions executable by the processor to:

randomly select a training signal from the set of training data;

identify a noise scenario associated with the selected training signal; and

determine a label for the selected training signal as a linear interpolation of a hard label for the training signal and a soft label from a teacher model trained using a subset of the training data associated with the identified noise scenario.

11. The system of claim 8 , wherein the memory stores instructions executable by the processor to:

input data based on an audio signal to the student model to obtain a transcript of speech recorded in the audio signal.

12. The system of claim 8 , wherein the a random scale factor is chosen from a range uniformly sampled in [−8 dB, −1 dB).

13. The system of claim 8 , wherein the message component is an audio signal recorded with a microphone near a desired audio source while the training signal is recorded with a microphone far from the desired audio source.

14. The system of claim 8 , wherein the memory stores instructions executable by the processor to:

apply feature extraction, including a log-mel filter bank, to the training signal and to the message component; and

subtract features of the message component from features of the training signal to obtain features of the noise component.

15. A non-transitory computer-readable storage medium, comprising executable instructions that, when executed by a processor, facilitate performance of operations, comprising:

training a model using audio recordings from noise scenarios in a set of training data;

decomposing a training signal from the set of training data into a message component and a noise component;

scaling the noise component by a random scale factor to obtain a scaled noise, wherein the random scale factor is a power with a base that is a constant and an exponent that includes a random variable;

adding the scaled noise to the message component to obtain a perturbed audio signal that is included in the set of training data;

training a first teacher model using a first subset of the set of training data associated with a first noise scenario of the noise scenarios;

training a second teacher model using a second subset of the set of training data associated with a second noise scenario of the noise scenarios; and

training a student model using soft labels output from the first teacher model and soft labels output from the second teacher model.

16. The non-transitory computer-readable storage medium of claim 15 , wherein training the student model using soft labels output from the second teacher model comprises:

determining a label for a training signal from the set of training data as a linear interpolation of a soft label from the second teacher model and a hard label for the training signal.

17. The non-transitory computer-readable storage medium of claim 15 , wherein the random scale factor is chosen from a range uniformly sampled in [−8 dB, −1 dB).

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2021
From: CHONG, DADING; LIU, ZHAOYI; PARTHASARATHY, VIJAY; SONG, XIAO
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 057041/0557 →
Continuity (1)
Related Publication 20230033768A1 · Feb 2, 2023
References Cited (18)
US 9495955B1 · Weber · 2016 [cited by examiner]
US 20190205748A1 · Fukuda · 2019 [cited by examiner]
US 20210224660A1 · Song · 2021 [cited by examiner]
EP 1424685A1 · 2004 [cited by examiner]
You Z, Su D, Yu D. Teach an all-rounder with experts in different domains. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) May 12, 2019 (pp. 6425-6429). IEEE. (Year:… [cited by examiner]
Fukuda T, Kurata G. Generalized knowledge distillation from an ensemble of specialized teachers leveraging unsupervised neural clustering. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signa… [cited by examiner]
Rebai I, BenAyed Y, Mahdi W, Lorre JP. Improving speech recognition using data augmentation and acoustic model fusion. Procedia Computer Science. Jan. 1, 2017;112:316-22. (Year: 2017). [cited by examiner]
Chen J, Wang Y, Wang D. Noise perturbation for supervised speech separation. Speech communication. Apr. 1, 2016;78:1-0. (Year: 2016). [cited by examiner]
Lincoln M, McCowan I, Vepa J, Maganti HK. The multi-channel Wall Street Journal audio visual corpus (MC-WSJ-AV): Specification and initial experiments. In IEEE Workshop on Automatic Speech Recognition and Understanding,… [cited by examiner]
Gao Y, Parcollet T, Lane N. Distilling Knowledge from Ensembles of Acoustic Models for Joint CTC-Attention End-to-End Speech Recognition. arXiv preprint arXiv:2005.09310. May 19, 2020. (Year: 2020). [cited by examiner]
Jaitly N, Hinton GE. Vocal tract length perturbation (VTLP) improves speech recognition. InProc. ICML Workshop on Deep Learning for Audio, Speech and Language Jun. 16, 2013 (vol. 117, p. 21). (Year: 2013). [cited by examiner]
Q. An, K. Bai, M. Zhang, Y. Yi and Y. Liu, “Deep Neural Network Based Speech Recognition Systems Under Noise Perturbations,” 2020 21st International Symposium on Quality Electronic Design (ISQED), Santa Clara, CA, USA, … [cited by examiner]
Semi-supervised end-to-end ASR via teacher-student learning with conditional posterior distribution, Zi-qiang Zhang et al., Interspeech 2020, Oct. 25-29, 20202, Shanghai, China, 5 pages. [cited by applicant]
Investigation of Data Augmentation Techniques for Disordered Speech Recognition, Mengzhe Geng et al., Interspeech 2020, Oct. 25-29, 2020, Shanghai, China, 5 pages. [cited by applicant]
Improving Noise Robustness of Automatic Speech Recognition Via Parallel Data and Teacher-Student Learning, Ladislav Mosner et al., Mar. 15, 2019, 5 pages. [cited by applicant]
Domain Adaptation Via Teacher-Student Learning for End-To-End Speech Recognition, Zhong Meng et al., Microsoft Corporation, Redmond, WA, USA, Jan. 6, 2020, 8 pages. [cited by applicant]
Audio Augmentation for Speech Recognition, Tom Ko et al., Johns Hopkins University, Baltimore, MD, 21218, USA, 4 pages, Year: 2015. [cited by applicant]
Improving speech recognition using data augmentation and acoustic model fusion, Ilyes Rebai et al., ScienceDirect, www.sciencedirect.com, Sep. 6-8, 2017, 7 pages. [cited by applicant]
Cited By (1)
US 12,609,111