IP Library › Granted Patent US 12,645,422
Granted Patent B2
US 12,645,422 · App. 18/376,019 · Granted Jun 2, 2026

Adaptive audio processing during video conferencing

Inventors: Yuhui Chen (San Jose, CA); Qiang Gao (Charlotte, NC); Zhaofeng Jia (Saratoga, CA); Shiwei Wang (Hefei, CN)
Assignee: Zoom Communications, Inc.
G06F3/165H04L12/1813
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,422
App. No.
18/376,019
Filed
Oct 3, 2023
Granted
Jun 2, 2026
Kind
B2
Examiner
ANWAH, OLISA
Art Unit
2692
USPC
348/14.01
Abstract

Techniques for adaptive audio processing during video conferencing are provided. In an example method, a client device joins a video conference hosted by a video conference provider, the video conference including a plurality of client devices. The client device receives, from an audio input device, an audio stream. The client device then processes, using an audio processing component, the audio stream. The client device determines, using a trained machine learning model, one or more characteristics of the audio stream. The client device then determines, based on the one or more characteristics of the audio stream, an audio configuration operation comprising one or more instructions. The client device executes the audio configuration operation. The client device then outputs the audio stream to the video conference provider.

Claims (88)

1 . A computer-implemented method, comprising:

joining a video conference hosted by a video conference provider, the video conference including a plurality of client devices;

receiving, from an audio input device, an audio stream;

determining, by a trained machine learning model, one or more characteristics of the audio stream;

determining, based on the one or more characteristics of the audio stream, an audio configuration operation comprising one or more instructions, the one or more instructions comprising application programming interface (“API”) commands for programmatically modifying one or more audio configurations of the client devices;

executing the audio configuration operation;

processing, by an audio processing component, the audio stream; and

outputting, to the video conference provider, the audio stream.

2 . The method of claim 1 , wherein the audio configuration operation includes instructions to modify at least one of the one or more characteristics of the audio stream.

3 . The method of claim 1 , wherein:

processing the audio stream comprises:

receiving, by the audio processing component, the audio stream;

executing, by the audio processing component, an audio processing operation, wherein the audio processing operation includes a modification of at least one of the one or more characteristics of the audio stream; and

the audio configuration operation determines the modification of at least one of the one or more characteristics of the audio stream.

4 . The method of claim 1 , wherein:

the audio processing component is a noise suppression component;

the audio configuration operation includes instructions to change a noise suppression mode; and

processing the audio stream comprises:

receiving, by the noise suppression component, the audio stream;

executing the instructions to change the noise suppression mode; and

executing, by the noise suppression component, a noise suppression operation, corresponding to the changed noise suppression mode.

5 . The method of claim 1 , wherein the one or more characteristics of the audio stream include one or more of speech quality, microphone quality, single-speaker, babble, music, or reverberation.

6 . The method of claim 1 , wherein:

the one or more characteristics of the audio stream include speech quality and microphone quality; and

the audio configuration operation includes instructions to perform real-time speech quality monitoring, comprising:

determining that at least one of the characteristics including speech quality or microphone quality is below a pre-determined threshold for a pre-determined period of time;

generating a notification comprising an indication of at least one of the speech quality or microphone quality; and

outputting the notification.

7 . The method of claim 1 , wherein:

the audio processing component is a noise suppression component;

the audio configuration operation includes instructions to change a noise suppression mode; and

determining, based on the one or more characteristics of the audio stream, the audio configuration operation comprises determining the instructions to change the noise suppression mode using an audio configuration operation logic table.

8 . The method of claim 1 , wherein:

the audio processing component is a noise suppression component;

the audio configuration operation includes instructions to change a noise suppression mode, wherein the noise suppression mode includes at least one of a high noise suppression mode, a medium noise suppression mode, or a low noise suppression mode; and

determining, based on the one or more characteristics of the audio stream, the audio configuration operation comprises determining the instructions to change the noise suppression mode, based on the one or more characteristics including at least microphone quality, single-speaker, babble, and music.

9 . The method of claim 1 , wherein the audio configuration operation comprises instructions to output a notification to a client device, wherein the notification includes a message to change an audio configuration based on the one or more characteristics.

10 . The method of claim 1 , wherein receiving, from the audio input device, the audio stream comprises:

segmenting the audio stream into one or more audio segments; and

for each audio segment, executing a signal processing instruction comprising at least one of short-time Fourier transform or conversion to a Mel spectrogram.

11 . A system comprising:

one or more computer systems, wherein the one or more computer systems are configured to perform processing comprising:

joining a video conference hosted by a video conference provider, the video conference including a plurality of client devices;

receiving, from an audio input device, an audio stream;

determining, by a trained machine learning model, one or more characteristics of the audio stream;

determining, based on the one or more characteristics of the audio stream, an audio configuration operation comprising one or more instructions, the one or more instructions comprising API commands for programmatically modifying one or more audio configurations of the client devices;

executing the audio configuration operation;

processing, by an audio processing component, the audio stream; and

outputting, to the video conference provider, the audio stream.

12 . The system of claim 11 , wherein the audio configuration operation includes instructions to modify at least one of the one or more characteristics of the audio stream.

13 . The system of claim 11 , wherein:

processing the audio stream comprises:

receiving, by the audio processing component, the audio stream;

executing, by the audio processing component, an audio processing operation, wherein the audio processing operation includes a modification of at least one of the one or more characteristics of the audio stream; and

the audio configuration operation determines the modification of at least one of the one or more characteristics of the audio stream.

14 . The system of claim 11 , wherein:

the audio processing component is a noise suppression component;

the audio configuration operation includes instructions to change a noise suppression mode; and

processing the audio stream comprises:

receiving, by the noise suppression component, the audio stream;

executing the instructions to change the noise suppression mode; and

executing, by the noise suppression component, a noise suppression operation, corresponding to the changed noise suppression mode.

15 . The system of claim 11 , wherein:

the one or more characteristics of the audio stream include one or more of speech quality, microphone quality, single-speaker, babble, music, or reverberation; and

the audio configuration operation includes instructions to perform real-time speech quality monitoring, comprising:

determining that at least one of the characteristics including speech quality or microphone quality is below a pre-determined threshold for a pre-determined period of time;

generating a notification comprising an indication of at least one of the speech quality or microphone quality; and

outputting the notification.

16 . A computer program product, comprising a computer program or instructions which, when executed by a processor, cause the processor to perform processing including:

joining a video conference hosted by a video conference provider, the video conference including a plurality of client devices;

receiving, from an audio input device, an audio stream;

determining, by a trained machine learning model, one or more characteristics of the audio stream;

determining, based on the one or more characteristics of the audio stream, an audio configuration operation comprising one or more instructions, the one or more instructions comprising API commands for programmatically modifying one or more audio configurations of the client devices;

executing the audio configuration operation;

processing, by an audio processing component, the audio stream; and

outputting, to the video conference provider, the audio stream.

17 . The computer program product of claim 16 , wherein:

the audio processing component is a noise suppression component;

the audio configuration operation includes instructions to change a noise suppression mode; and

determining, based on the one or more characteristics of the audio stream, the audio configuration operation comprises determining the instructions to change the noise suppression mode using an audio configuration operation logic table.

18 . The computer program product of claim 16 , wherein:

the audio processing component is a noise suppression component;

the audio configuration operation includes instructions to change a noise suppression mode, wherein the noise suppression mode includes at least one of a high noise suppression mode, a medium noise suppression mode, or a low noise suppression mode; and

determining, based on the one or more characteristics of the audio stream, the audio configuration operation comprises determining the instructions to change the noise suppression mode, based on the one or more characteristics including at least microphone quality, single-speaker, babble, and music.

19 . The computer program product of claim 16 , wherein the audio configuration operation comprises instructions to output a notification to a client device, wherein the notification includes a message to change an audio configuration based on the one or more characteristics.

20 . The computer program product of claim 16 , wherein receiving, from the audio input device, the audio stream comprises:

segmenting the audio stream into one or more audio segments; and

for each audio segment, executing a signal processing instruction comprising at least one of short-time Fourier transform or conversion to a Mel spectrogram.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2026
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 076007/0346 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 3, 2023
From: CHEN, YUHUI; GAO, QIANG; JIA, ZHAOFENG; WANG, SHIWEI
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 065103/0459 →
Priority Claims (1)
WO PCT/CN2023/100750 · Jun 16, 2023 · international
Continuity (1)
Related Publication 20240419391A1 · Dec 19, 2024
References Cited (7)
US 10446170B1 · Chen · 2019 [cited by examiner]
US 11621016B2 · Deng · 2023 [cited by examiner]
US 20160005422A1 · Zad Issa et al. · 2016 [cited by applicant]
US 20210360349A1 · Nyayate · 2021 [cited by examiner]
US 20220051652A1 · Winsvold et al. · 2022 [cited by applicant]
US 20240127839A1 · Clark · 2024 [cited by examiner]
EP International Search Report and Written Opinion for PCT/CN2023/100750 mailed Dec. 1, 2023. [cited by applicant]