Intelligent Muting Of Participant Audio In Communication Sessions
An input audio signal associated with a communication session is received. A determination is made that a participant associated with the input audio signal is not audibly speaking within the input audio signal. In response to determining that the participant is not audibly speaking, an audio feed corresponding to the input audio signal is muted by rendering the audio feed not audible to at least one other participant device connected to the communication session.
1 . A method, comprising:
receiving an input audio signal associated with a communication session;
determining that a participant associated with the input audio signal is not audibly speaking within the input audio signal; and
in response to determining that the participant is not audibly speaking, muting an audio feed corresponding to the input audio signal by rendering the audio feed not audible to at least one other participant device connected to the communication session.
2 . The method of claim 1 , further comprising:
analyzing the input audio signal for audible speech from the participant;
detecting that the participant is audibly speaking within the input audio signal; and
unmuting the audio feed transmitted to the at least one other participant device.
3 . The method of claim 2 , wherein detecting that the participant is audibly speaking within the input audio signal comprises recognizing that the participant has uttered a prespecified passphrase.
4 . The method of claim 2 , further comprising:
writing or overwriting a recording buffer with content of the input audio signal, wherein unmuting the audio feed transmitted to the at least one other participant device comprises:
playing back a portion of the recording buffer comprising vocal speech of the participant.
5 . The method of claim 2 , wherein detecting that the participant is audibly speaking within the input audio signal comprises:
recognizing that the participant has performed a prespecified non-verbal gesture within a video feed of the participant.
6 . The method of claim 1 , further comprising:
in response to muting the audio feed, sending a first alert to the participant at a client device.
7 . The method of claim 1 , further comprising:
detecting audible speech that is not from the participant within the input audio signal; and
muting the audio feed transmitted to the at least one other participant device.
8 . The method of claim 1 , further comprising:
detecting audible speech from the participant and one or more additional participants concurrently within the communication session;
determining a speaking order of concurrently speaking participants; and
based on the speaking order, muting the audio feed of the participant.
9 . A system, comprising:
one or more memories; and
one or more processors, the one or more processors configured to execute instructions stored in the one or more memories to:
receive an input audio signal associated with a communication session;
determine that a participant associated with the input audio signal is not audibly speaking within the input audio signal; and
in response to determining that the participant is not audibly speaking, mute an audio feed corresponding to the input audio signal by rendering the audio feed not audible to at least one other participant device connected to the communication session.
10 . The system of claim 9 , the one or more processors further configured to execute instructions in the one or more memories to:
analyze the input audio signal for audible speech from the participant; and
in response to detecting that the participant is audibly speaking within the input audio signal, unmute the audio feed transmitted to the at least one other participant device.
11 . The system of claim 10 , wherein, to determine that the participant is not audibly speaking within the input audio signal, the one or more processors configured to execute instructions stored in the one or more memories to:
periodically analyze waveforms of the input audio signal to detect an absence representative of vocal speech of the participant, wherein to periodically analyze waveforms the one or more processors configured to execute instructions stored in the one or more memories to:
determine a representative amplitude of the vocal speech; and
detect, within the waveforms, a decrease in amplitude proportional to the representative amplitude of the vocal speech.
12 . The system of claim 10 , wherein, to detect that the participant is audibly speaking within the input audio signal, the one or more processors configured to execute instructions stored in the one or more memories to:
recognize that the participant has uttered a custom passphrase selected by the participant.
13 . The system of claim 9 , wherein, to determine that the participant is not audibly speaking, the one or more processors configured to execute instructions stored in the one or more memories to:
determine that the participant is not audibly speaking using an artificial intelligence model that is trained to recognize vocal speech of the participant within an audio signal.
14 . The system of claim 9 , the one or more processors further configured to execute instructions in the one or more memories to:
in response to muting the audio feed, send a first alert to the participant; and
in response to unmuting the audio feed, send a second alert different from the first alert to the participant.
15 . The system of claim 14 , wherein the first alert and the second alert each comprise one or more of: a vibration alert, an audio alert, and a visual alert.
16 . The system of claim 9 , wherein, to determine that the participant is not audibly speaking, the one or more processors configured to execute instructions stored in the one or more memories to:
extract audio features from the input audio signal; and
use the audio features as input for a machine learning model that outputs a classification prediction regarding whether a voice of the participant is audibly present in the input audio signal.
17 . A non-transitory computer-readable storage medium, comprising executable instructions that, when executed by a processor, perform operations, comprising:
receiving an input audio signal associated with a communication session;
determining that a participant associated with the input audio signal is not audibly speaking within the input audio signal; and
in response to determining that the participant is not audibly speaking, muting an audio feed corresponding to the input audio signal by rendering the audio feed not audible to at least one other participant device connected to the communication session.
18 . The non-transitory computer-readable storage medium of claim 17 , wherein determining that the participant is not audibly speaking comprises:
extracting audio features from the input audio signal; and
providing the audio features to a trained artificial intelligence model, wherein the audio features comprise at least one of Mel-frequency cepstral coefficients or spectral peaks, and wherein the trained artificial intelligence model outputs a probability label indicative of whether the participant is audibly speaking within the input audio signal.
19 . The non-transitory computer-readable storage medium of claim 17 , wherein muting the audio feed comprises using an artificial intelligence based silencing technique trained on voice samples of the participant and background noise.
20 . The non-transitory computer-readable storage medium of claim 17 , wherein determining that the participant is not audibly speaking comprises:
detecting audible speech that is not from the participant within the input audio signal.