EXTRANEOUS VOICE REMOVAL FROM AUDIO IN A COMMUNICATION SESSION
The technology disclosed herein enables removal of extraneous voices from audio in a communication session. In a particular embodiment, a method includes receiving audio captured from an endpoint operated by a user on a communication session. The method further includes identifying an extraneous voice in the audio, wherein the voice is from a person other than the user, and removing the extraneous voice from the audio. After removing the extraneous voice, the method includes transmitting the audio to another endpoint on the communication session.
1 . A method comprising:
receiving audio comprising a signal representing sound captured from an endpoint operated by a user on a communication session;
identifying an extraneous voice in the audio, wherein the voice is from a person other than the user;
removing the extraneous voice from the audio; and
after removing the extraneous voice, transmitting the audio to another endpoint on the communication session.
2 . The method of claim 1 , wherein identifying and removing the extraneous voice comprise:
inputting the audio into a machine learning algorithm, wherein the machine learning algorithm is trained to recognize a user voice of the user and wherein the machine learning algorithm outputs the audio with the extraneous voice removed.
3 . The method of claim 2 , comprising:
training the machine learning algorithm using one or more samples of the user voice.
4 . The method of claim 3 , wherein training the machine learning algorithm includes:
in response to the user initiating the communication session, requesting the samples from the user.
5 . The method of claim 3 , wherein training the machine learning algorithm comprises:
training the machine learning algorithm using one or more extraneous voice samples that were not intended for transmittal.
6 . The method of claim 3 , wherein the machine learning algorithm generates a confidence score for the extraneous voice and removes the extraneous voice upon determining that the confidence score satisfies a threshold level of confidence.
7 . The method of claim 6 , wherein the machine learning algorithm considers intensity of the extraneous voice and/or a language spoken by the extraneous voice when generating the confidence score.
8 . The method of claim 1 , comprising:
notifying the user that the extraneous voice has been identified;
wherein removing the extraneous voice is performed in response to determining that the user has granted permission for removal of the extraneous voice.
9 . The method of claim 1 , wherein identifying the extraneous voice comprises:
isolating the extraneous voice from one or more other voices in the audio.
10 . The method of claim 1 , wherein the extraneous voice is not a voice included in a whitelist of voices.
11 . An apparatus comprising:
one or more computer readable storage media;
a processing system operatively coupled with the one or more computer readable storage media; and
program instructions stored on the one or more computer readable storage media that, when read and executed by the processing system, direct the processing system to:
receive audio comprising a signal representing sound captured from an endpoint operated by a user on a communication session;
identify an extraneous voice in the audio, wherein the voice is from a person other than the user;
remove the extraneous voice from the audio; and
after removing the extraneous voice, transmit the audio to another endpoint on the communication session.
12 . The apparatus of claim 11 , wherein to identify and remove the extraneous voice, the program instructions direct the processing system to:
input the audio into a machine learning algorithm, wherein the machine learning algorithm is trained to recognize a user voice of the user and wherein the machine learning algorithm outputs the audio with the extraneous voice removed.
13 . The apparatus of claim 12 , wherein the program instructions direct the processing system to:
train the machine learning algorithm using one or more samples of the user voice.
14 . The apparatus of claim 13 , wherein to train the machine learning algorithm, the program instructions direct the processing system to:
in response to the user initiating the communication session, request the samples from the user.
15 . The apparatus of claim 13 , wherein to train the machine learning algorithm, the program instructions direct the processing system to:
train the machine learning algorithm using one or more extraneous voice samples that were not intended for transmittal.
16 . The apparatus of claim 13 , wherein the machine learning algorithm generates a confidence score for the extraneous voice and removes the extraneous voice upon determining that the confidence score satisfies a threshold level of confidence.
17 . The apparatus of claim 16 , wherein the machine learning algorithm considers intensity of the extraneous voice and/or a language spoken by the extraneous voice when generating the confidence score.
18 . The apparatus of claim 11 , wherein the program instructions direct the processing system to:
notify the user that the extraneous voice has been identified;
wherein removal of the extraneous voice is performed in response to determining that the user has granted permission for removal of the extraneous voice.
19 . The apparatus of claim 11 , wherein identifying the extraneous voice the program instructions direct the processing system to:
isolating the extraneous voice from one or more other voices in the audio.
20 . One or more computer readable storage media having program instructions stored thereon that, when read and executed by a processing system, direct the processing system to:
receive audio comprising a signal representing sound captured from an endpoint operated by a user on a communication session;
identify an extraneous voice in the audio, wherein the voice is from a person other than the user;
remove the extraneous voice from the audio; and
after removing the extraneous voice, transmit the audio to another endpoint on the communication session.