IP Library Granted Patent US 11,621,016
Granted Patent B2
US 11,621,016 · App. 17/390,915 · Granted Apr 4, 2023

Intelligent noise suppression for audio signals within a communication platform

Inventors: Jiachuan Deng (Mountain View, CA); Qiyong Liu (Santa Clara, CA); Chuanfei Wang (Zhejiang, CN); Xiuyu Xu (Zhejiang, CN)
Assignee: Zoom Video Communications, Inc.
G10L21/0232G06N3/08G06N7/005G10L25/18G10L25/30G10L25/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,621,016
App. No.
17/390,915
Granted
Apr 4, 2023
Kind
B2
Abstract

Methods and systems provide users of a communication platform with intelligent, real-time noise suppression for audio signals broadcasted in a communication session. The system receives an input audio signal from an audio capture device; processes the input audio signal to provide a second version of the audio signal with noise suppression based on DSP techniques; transmits the second version of the audio signal to a communication platform for real-time streaming; classifies, via a machine learning algorithm, whether the second version of the audio signal contains noise beyond a noise threshold; based on a classification that the second version of the audio signal contains noise beyond the noise threshold, processes the second version of the audio signal to provide a third version of the audio signal with noise suppression based on AI techniques; and transmits the third version of the audio signal to the communication platform.

Claims (59)

1. A communication system comprising one or more processors configured to perform the operations of:

receiving an input audio signal from an audio capture device;

processing the input audio signal to provide a second version of the audio signal with noise suppression based on digital signal processing (DSP) techniques;

transmitting the second version of the audio signal to a communication platform for real-time streaming;

classifying, via a machine learning algorithm, whether the second version of the audio signal contains noise beyond a noise threshold;

based on a classification that the second version of the audio signal contains noise beyond the noise threshold, processing the second version of the audio signal to provide a third version of the audio signal with noise suppression based on artificial intelligence (AI) techniques; and

transmitting the third version of the audio signal to the communication platform.

2. The system of claim 1 , wherein the second version of the audio signal provides suppression of stationary noises, and wherein the third version of the audio signal provides suppression of both stationary and non-stationary noises.

3. The system of claim 1 , wherein the noise threshold is determined by the machine learning algorithm.

4. The system of claim 1 , wherein classifying whether the second version of the audio signal contains noise beyond a noise threshold comprises:

extracting a plurality of audio features from the input audio signal, wherein the input audio signal is a raw waveform;

transmitting the audio features to a neural network; and

analyzing the audio features via the neural network to provide a probability of whether the second version of the audio signal contains noise beyond the noise threshold.

5. The system of claim 4 , wherein the neural network comprises at least one of a convolutional neural network (CNN) and a multilayer perceptron (MLP).

6. The system of claim 4 , further comprising:

generating a spectrogram based on the extracted audio features,

wherein transmitting the audio features to the neural network comprises transmitting the spectrogram to the neural network, and

wherein analyzing the audio features via the neural network comprises analyzing the spectrogram.

7. The system of claim 1 , wherein classifying whether the second version of the audio signal contains noise beyond a noise threshold comprises:

generating an output label comprising the classification after a predefined time interval has expired; and

storing the output label within a buffer, wherein the buffer contains a plurality of output labels generated within a predefined window of time.

8. The system of claim 7 , wherein classifying whether the second version of the audio signal contains noise beyond a noise threshold further comprises:

determining that the predefined window of time has expired;

generating a confidence score for output labels stored within the buffer; and

based on the confidence score, determining whether the second version of the audio signal meets or exceeds the noise threshold.

9. The system of claim 1 , wherein receiving the input audio signal and processing the input audio signal are performed by a client device associated with a user.

10. The system of claim 1 , wherein transmitting the third version of the audio signal to the communication platform is performed in real time or substantially real time after transmitting the second version of the audio signal to the communication platform for real-time streaming.

11. The system of claim 1 , further comprising:

performing one or more additional DSP techniques on the third version of the audio signal.

12. The system of claim 1 , further comprising:

providing real-time streaming at the communication platform using the third version of the audio signal.

13. The system of claim 1 , wherein classifying whether the second version of the audio signal contains noise beyond a noise threshold comprises performing one or more feature-based classification techniques.

14. A method for providing intelligent noise suppression for an audio signal within a communication platform, comprising:

receiving an input audio signal from an audio capture device;

processing the input audio signal to provide a second version of the audio signal with noise suppression based on digital signal processing (DSP) techniques;

transmitting the second version of the audio signal to a communication platform for real-time streaming;

classifying, via a machine learning algorithm, whether the second version of the audio signal contains noise beyond a noise threshold;

based on a classification that the second version of the audio signal contains noise beyond the noise threshold, processing the second version of the audio signal to provide a third version of the audio signal with noise suppression based on artificial intelligence (AI) techniques; and

transmitting the third version of the audio signal to the communication platform.

15. The method of claim 14 , wherein the second version of the audio signal provides suppression of stationary noises, and wherein the third version of the audio signal provides suppression of both stationary and non-stationary noises.

16. The method of claim 14 , wherein the noise threshold is determined by the machine learning algorithm.

17. The method of claim 14 , wherein classifying whether the second version of the audio signal contains noise beyond a noise threshold comprises:

extracting a plurality of audio features from the input audio signal, wherein the input audio signal is a raw waveform;

transmitting the audio features to a neural network; and

analyzing the audio features via the neural network to provide a probability of whether the second version of the audio signal contains noise beyond the noise threshold.

18. The method of claim 17 , further comprising:

generating a spectrogram based on the extracted audio features,

wherein transmitting the audio features to the neural network comprises transmitting the spectrogram to the neural network, and

wherein analyzing the audio features via the neural network comprises analyzing the spectrogram.

19. The method of claim 14 , wherein classifying whether the second version of the audio signal contains noise beyond a noise threshold comprises:

generating an output label comprising the classification after a predefined time interval has expired; and

storing the output label within a buffer, wherein the buffer contains a plurality of output labels generated within a predefined window of time.

20. A non-transitory computer-readable medium containing instructions for dynamically altering notification preferences within a communication platform, comprising:

instructions for receiving an input audio signal from an audio capture device;

instructions for processing the input audio signal to provide a second version of the audio signal with noise suppression based on digital signal processing (DSP) techniques;

instructions for transmitting the second version of the audio signal to a communication platform for real-time streaming;

instructions for classifying, via a machine learning algorithm, whether the second version of the audio signal contains noise beyond a noise threshold;

based on a classification that the second version of the audio signal contains noise beyond the noise threshold, instructions for processing the second version of the audio signal to provide a third version of the audio signal with noise suppression based on artificial intelligence (AI) techniques; and

instructions for transmitting the third version of the audio signal to the communication platform.

Assignments (2)
CHANGE OF NAME Recorded Sep 25, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 072919/0122 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2023
From: DENG, JIACHUAN; LIU, QIYONG; XU, XIUYU; WANG, CHUANFEI
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 062690/0063 →
Continuity (1)
Related Publication 20230032785A1 · Feb 2, 2023
Cited By (1)
US 12,645,422