Double talk detection using capture up-sampling
A method of double talk detection includes using up-sampling. Audio signals received from the far end are up-sampled prior to output by the loudspeaker at the near end. The microphone at the near end captures audio at the up-sampled rate, and the audio output by the loudspeaker is detectable due to having no energy in the up-sampled frequency bands. The double talk detector uses this information to generate a signal for suppressing the echo of the far end audio from the captured audio signal that is transmitted to the far end.
1 . A computer-implemented method of audio processing, the method comprising:
receiving a first audio signal, wherein the first audio signal has a first sampling frequency;
up-sampling the first audio signal to generate a second audio signal, wherein the second audio signal has a second sampling frequency that is greater than the first sampling frequency;
outputting, by a loudspeaker, a loudspeaker output corresponding to the second audio signal;
capturing, by a microphone, a third audio signal, wherein the third audio signal is sampled at the second sampling frequency;
determining a signal power of the third audio signal; and
detecting double talk when there is signal power of the third audio signal determined in a frequency band having frequencies all greater than half the first sampling frequency.
2 . The method of claim 1 , further comprising:
selectively generating a control signal when the double talk is detected; and
performing echo management on the third audio signal according to the control signal.
3 . The method of claim 2 , wherein performing echo management includes:
performing echo cancellation on the third audio signal according to the control signal, wherein the echo cancellation performs a linear attenuation on the third audio signal.
4 . The method of claim 2 , wherein performing echo management includes:
performing echo suppression on the third audio signal according to the control signal, wherein the echo suppression performs a non-linear attenuation on particular frequency bands of the third audio signal.
5 . The method of claim 1 , wherein the third audio signal includes local audio and the loudspeaker output, wherein the local audio corresponds to audio other than the loudspeaker output, and wherein the local audio is not outputted by the loudspeaker and is captured by the microphone.
6 . The method of claim 1 , wherein the first sampling frequency is 8 kHz, and wherein the second sampling frequency is at least 16 kHz.
7 . The method of claim 1 , further comprising:
down-sampling the third audio signal to generate a fourth audio signal, wherein the fourth audio signal has a third sampling frequency that is less than the second sampling frequency; and
transmitting the fourth audio signal to a far end device.
8 . The method of claim 7 , wherein the third sampling frequency and the first sampling frequency are the same sampling frequency.
9 . The method of claim 1 , wherein determining the signal power of the third audio signal and detecting the double talk includes:
measuring the signal power of the third audio signal in the frequency band greater than the first sampling frequency;
tracking a background noise power of the third audio signal in the frequency band greater than the first sampling frequency; and
detecting the double talk as a result of comparing the signal power of the third audio signal in the frequency band having frequencies all greater than half the first sampling frequency and the background noise power of the third audio signal in the frequency band having frequencies all greater than half the first sampling frequency.
10 . The method of claim 1 , wherein determining the signal power of the third audio signal and detecting the double talk includes:
measuring the signal power of the third audio signal in the frequency band greater than the first sampling frequency;
tracking a background noise power of the third audio signal in the frequency band greater than the first sampling frequency;
measuring a distortion power of the first audio signal; and
detecting the double talk based on the signal power of the third audio signal in the frequency band having frequencies all greater than half the first sampling frequency, the background noise power of the third audio signal in the frequency band having frequencies all greater than half the first sampling frequency, and the distortion power of the first audio signal.
11 . The method of claim 10 , wherein measuring the distortion power of the first audio signal includes:
generating a filtered signal by performing band pass filtering on the first audio signal;
measuring a signal power of the filtered signal; and
determining the distortion power by performing non-linear regulation on the signal power of the filtered signal.
12 . A non-transitory computer readable medium storing a computer program that, when executed by a processor, controls an apparatus to execute processing including the method of claim 1 .
13 . An apparatus for audio processing, the apparatus comprising:
a loudspeaker;
a microphone; and
a processor;
wherein the processor is configured to control the apparatus to receive a first audio signal, wherein the first audio signal has a first sampling frequency;
wherein the processor is configured to control the apparatus to up-sample the first audio signal to generate a second audio signal, wherein the second audio signal has a second sampling frequency that is greater than the first sampling frequency;
wherein the processor is configured to control the apparatus to output, by the loudspeaker, a loudspeaker output corresponding to the second audio signal;
wherein the processor is configured to control the apparatus to capture, by the microphone, a third audio signal, wherein the third audio signal is sampled at the second sampling frequency;
wherein the processor is configured to control the apparatus to determine a signal power of the third audio signal; and
wherein the processor is configured to control the apparatus to detect double talk when there is signal power of the third audio signal determined in a frequency band having frequencies all greater than half the first sampling frequency.
14 . The apparatus of claim 13 , wherein the processor is configured to control the apparatus to selectively generate a control signal when the double talk is detected; and
wherein the processor is configured to control the apparatus to perform echo management on the third audio signal according to the control signal.
15 . The apparatus of claim 14 , wherein controlling the apparatus to perform echo management includes:
controlling the apparatus to perform echo cancellation on the third audio signal according to the control signal, wherein the echo cancellation performs a linear attenuation on the third audio signal.
16 . The apparatus of claim 14 , wherein controlling the apparatus to perform echo management includes:
controlling the apparatus to perform echo suppression on the third audio signal according to the control signal, wherein the echo suppression performs a non-linear attenuation on particular frequency bands of the third audio signal.
17 . The apparatus of claim 13 , wherein the processor is configured to control the apparatus to down-sample the third audio signal to generate a fourth audio signal, wherein the fourth audio signal has a third sampling frequency that is less than the second sampling frequency; and
wherein the processor is configured to control the apparatus to transmit the fourth audio signal to a far end device.
18 . The apparatus of claim 13 , wherein controlling the apparatus to determine the signal power of the third audio signal and to detect the double talk includes:
controlling the apparatus to measure the signal power of the third audio signal in the frequency band greater than the first sampling frequency;
controlling the apparatus to track a background noise power of the third audio signal in the frequency band greater than the first sampling frequency; and
controlling the apparatus to detect the double talk as a result of comparing the signal power of the third audio signal in the frequency band having frequencies all greater than half the first sampling frequency and the background noise power of the third audio signal in the frequency band having frequencies all greater than half the first sampling frequency.
19 . The apparatus of claim 13 , wherein controlling the apparatus to determine the signal power of the third audio signal and to detect the double talk includes:
controlling the apparatus to measure the signal power of the third audio signal in the frequency band greater than the first sampling frequency;
controlling the apparatus to track a background noise power of the third audio signal in the frequency band greater than the first sampling frequency;
controlling the apparatus to measure a distortion power of the first audio signal; and
controlling the apparatus to detect the double talk based on the signal power of the third audio signal in the frequency band having frequencies all greater than the half first sampling frequency, the background noise power of the third audio signal in the frequency band having frequencies all greater than half the first sampling frequency, and the distortion power of the first audio signal.
20 . The apparatus of claim 19 , wherein controlling the apparatus to measure the distortion power of the first audio signal includes:
controlling the apparatus to generate a filtered signal by performing band pass filtering on the first audio signal;
controlling the apparatus to measure a signal power of the filtered signal; and
controlling the apparatus to determine the distortion power by performing non-linear regulation on the signal power of the filtered signal.