IP Library Granted Patent US 12,443,389
Granted Patent B1
US 12,443,389 · App. 18/316,055 · Granted Oct 14, 2025

Intelligent muting and unmuting of an audio feed within a communication session

Inventor: Thanh Le Nguyen (Belle Chasse, LA)
Assignee: Zoom Communications, Inc.
G06F3/165G10L17/04G10L17/24G10L25/78H04L65/403G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,443,389
App. No.
18/316,055
Granted
Oct 14, 2025
Kind
B1
Abstract

An input audio signal is received from a client device. A connection to a communication session with a plurality of participants is maintained for the client device. An audio feed corresponding to the input audio signal is transmitted to a plurality of other participant devices. While the communication session is in progress and the connection to the communication session is maintained by the client device, the following are periodically performed. An absence of audible speech originating at the client device within the input audio signal is detected. The audio feed transmitted to the plurality of other participant devices is muted. The input audio signal is analyzed for audible speech. A presence of audible speech within the input audio signal is detected. The audio feed transmitted to the plurality of other participant devices is unmuted.

Claims (71)

1. A method, comprising:

receiving an input audio signal from a client device;

maintaining, for the client device, a connection to a communication session with a plurality of participants, wherein an audio feed corresponding to the input audio signal is transmitted to a plurality of other participant devices; and

while the communication session is in progress and the connection to the communication session is maintained by the client device, periodically performing:

detecting an absence of audible speech originating at the client device within the input audio signal;

in response to detecting the absence of the audible speech within the input audio signal, muting the audio feed transmitted to the plurality of other participant devices;

analyzing the input audio signal for audible speech;

detecting audible speech of the participant within the input audio signal; and

in response to detecting the audible speech within the input audio signal, unmuting the audio feed transmitted to the plurality of other participant devices.

2. The method of claim 1 , wherein detecting the absence of audible speech originating at the client device within the input audio signal comprises:

recognizing an utterance of a prespecified passphrase.

3. The method of claim 2 , wherein the prespecified passphrase is a custom passphrase selected by a participant associated with the client device.

4. The method of claim 1 , wherein detecting the absence of audible speech originating at the client device within the input audio signal comprises:

periodically analyzing waveforms of the input audio signal to detect an absence of an expected vocal speech.

5. The method of claim 4 , wherein periodically analyzing the waveforms of the input audio signal to detect the absence of the expected vocal speech comprises:

determining a representative amplitude of an expected vocal speech, and

detecting, within the waveforms, a decrease in amplitude proportional to the representative amplitude of the expected vocal speech.

6. The method of claim 1 , wherein detecting the audible speech of the participant or the absence of the audible speech within the input audio signal is performed by an artificial intelligence (AI) model trained to recognize an expected vocal speech.

7. The method of claim 1 , further comprising:

writing or overwriting a recording buffer with content from the input audio signal,

wherein unmuting the audio feed transmitted to the plurality of other participant devices comprises:

playing back a portion of the recording buffer.

8. The method of claim 1 , further comprising:

in response to muting the audio feed transmitted to the plurality of other participant devices, sending a first alert to a participant associated with the client device; and

in response to unmuting the audio feed transmitted to the plurality of other participant devices, sending a second alert different from the first alert to the participant.

9. The method of claim 8 , wherein the first alert and the second alert each comprises one or more of a vibration alert, an audio alert, or a visual alert.

10. The method of claim 1 , further comprising:

detecting audible that is not from the participant within the input audio signal; and

muting the audio feed transmitted to the plurality of other participant devices.

11. The method of claim 1 , further comprising:

upon detecting the audible speech of the participant within the input audio signal, determining that the audio feed is independently muted via the client device and via one or more input devices, wherein unmuting the audio feed transmitted to the plurality of other participant devices comprises:

unmuting the audio feed at the client device and the one or more input devices.

12. The method of claim 1 , wherein detecting the audible speech of the participant within the input audio signal comprises:

recognizing a performance of a prespecified non-verbal gesture within a video feed originating at the client device.

13. The method of claim 1 , wherein detecting the audible speech of the participant within the input audio signal comprises:

recognizing a performance a prespecified non-verbal utterance.

14. The method of claim 1 , further comprising:

detecting audible speech from the client device and one or more additional participant devices concurrently within the communication session;

determining a speaking order; and

muting, based on the speaking order, the audio feed transmitted to the plurality of other participant devices.

15. A communication system comprising one or more processors configured to perform instructions to:

receive an input audio signal from a client device;

maintain, for the client device, a connection to a communication session with a plurality of participants, wherein an audio feed corresponding to the input audio signal is transmitted to a plurality of other participant devices; and

while the communication session is in progress and the connection to the communication session is maintained by the client device, periodically perform instructions to:

detect an absence of audible speech originating at the client device within the input audio signal;

in response to detecting the absence of the audible speech within the input audio signal, mute the audio feed transmitted to the plurality of other participant devices;

analyze the input audio signal for audible speech;

detect audible speech of the participant within the input audio signal; and

in response to detecting the audible speech within the input audio signal, unmute the audio feed transmitted to the plurality of other participant devices.

16. The communication system of claim 15 , wherein the instructions to detect the audible speech of the participant or the absence of the audible speech within the input audio signal are performed by an artificial intelligence (AI) model trained to recognize an expected vocal speech.

17. The communication system of claim 15 , wherein the one or more processors are further configured to execute instructions to:

write or overwrite a recording buffer with content from the input audio signal,

wherein to unmute the audio feed transmitted to the plurality of other participant devices comprises instructions to:

play back a portion of the recording buffer.

18. The communication system of claim 15 , wherein the one or more processors are further configured to execute instructions to:

upon detecting the audible speech of the participant within the input audio signal, determine that the audio feed is independently muted via the client device and one or more input devices,

wherein to unmute the audio feed transmitted to the plurality of other participant devices comprises to:

unmute the audio feed at the client device and the one or more input devices.

19. The communication system of claim 15 , wherein the one or more processors are further configured to execute instructions to:

detect audible speech from the client device and one or more additional participant devices concurrently within the communication session;

determine a speaking order; and

mute, based on the speaking order, the audio feed transmitted to the plurality of other participant devices.

20. A non-transitory computer-readable storage medium, comprising executable instructions that, when executed by a processor, facilitate performance of operations, comprising:

receiving an input audio signal from a client device;

maintaining, for the client device, a connection to a communication session with a plurality of participants, wherein an audio feed corresponding to the input audio signal is transmitted to a plurality of other participant devices; and

while the communication session is in progress and the connection to the communication session is maintained by the client device, periodically performing:

detecting an absence of audible speech originating at the client device within the input audio signal;

in response to detecting the absence of the audible speech within the input audio signal, muting the audio feed transmitted to the plurality of other participant devices;

analyzing the input audio signal for audible speech;

detecting audible speech of the participant within the input audio signal; and

in response to detecting the audible speech within the input audio signal, unmuting the audio feed transmitted to the plurality of other participant devices.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2023
From: NGUYEN, THANH LE
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 063617/0748 →
Continuity (1)
Continuation 17573454 · Jan 11, 2022
References Cited (35)
US 11176923B1 · Vendrow · 2021 [cited by examiner]
US 11297281B1 · Agrawal · 2022 [cited by examiner]
US 11489895B1 · Fardig · 2022 [cited by examiner]
US 11620041B1 · Boucheron · 2023 [cited by examiner]
US 11812194B1 · Vandyke · 2023 [cited by examiner]
US 20180124458A1 · Knox · 2018 [cited by examiner]
US 20180124459A1 · Knox · 2018 [cited by examiner]
US 20180349086A1 · Chakra · 2018 [cited by examiner]
US 20200103963A1 · Kelly · 2020 [cited by examiner]
US 20200110572A1 · Lenke · 2020 [cited by examiner]
US 20200274911A1 · Gargaro · 2020 [cited by examiner]
US 20200326846A1 · Leong · 2020 [cited by examiner]
US 20210051035A1 · Atkins · 2021 [cited by examiner]
US 20210051036A1 · Atkins · 2021 [cited by examiner]
US 20210051037A1 · Atkins · 2021 [cited by examiner]
US 20210383824A1 · Condorelli · 2021 [cited by examiner]
US 20210392175A1 · Gronau · 2021 [cited by examiner]
US 20210392231A1 · Gronau · 2021 [cited by examiner]
US 20210399911A1 · Jorasch · 2021 [cited by examiner]
US 20210400142A1 · Jorasch · 2021 [cited by examiner]
US 20220036708A1 · Rey · 2022 [cited by examiner]
US 20220051412A1 · Gronau · 2022 [cited by examiner]
US 20220051652A1 · Winsvold · 2022 [cited by examiner]
US 20220139383A1 · Rose · 2022 [cited by examiner]
US 20220182578A1 · Sircar · 2022 [cited by examiner]
US 20220198388A1 · Simpson · 2022 [cited by examiner]
US 20220345501A1 · Li · 2022 [cited by examiner]
US 20220377117A1 · Xi · 2022 [cited by examiner]
US 20220413794A1 · Qiao · 2022 [cited by examiner]
US 20230007056A1 · Cox · 2023 [cited by examiner]
US 20230038109A1 · Braganza · 2023 [cited by examiner]
US 20230077283A1 · Mehta · 2023 [cited by examiner]
US 20230090613A1 · Covell · 2023 [cited by examiner]
US 20230133539A1 · Sieracki · 2023 [cited by examiner]
US 20250088795A1 · Graham · 2025 [cited by examiner]