Clock skew robust acoustic echo cancellation
Far-end audio samples may be received corresponding to far-end audio that is output from one or more audio output components. Near-end audio samples may be received corresponding to near-end audio that is captured by one or more audio input components. A plurality of acoustic path estimates and a plurality of clock skew estimates may be calculated in an alternating order, using a state-space model, based at least in part on the far-end audio samples and the near-end audio samples. A first acoustic path estimate and a first clock skew estimate may be used to calculate a second acoustic path estimate. A first portion of the far-end audio may be filtered with the second acoustic path estimate to generate a replica of echo in the first portion of the far-end audio. The replica of the echo may be removed from a corresponding second portion of the near-end audio.
1 . A computing system comprising:
one or more processors; and
one or more memories having stored therein instructions that, upon execution by the one or more processors, cause the computing system to perform computing operations comprising:
receiving far-end audio samples corresponding to far-end audio that is output from one or more audio output components at a near-end location, wherein the far-end audio is captured at a far-end location and transmitted to the near-end location;
receiving near-end audio samples corresponding to near-end audio that is captured by one or more audio input components at the near-end location;
calculating, using a state-space model, based at least in part on the far-end audio samples and the near-end audio samples, a plurality of acoustic path estimates and a plurality of clock skew estimates, wherein the plurality of acoustic path estimates approximate an acoustic path between the one or more audio output components and the one or more audio input components, and wherein the plurality of clock skew estimates approximate a clock skew caused by a difference between a far-end sampling rate associated with the one or more audio output components and a near-end sampling rate associated with the one or more audio input components, wherein the plurality of acoustic path estimates and the plurality of clock skew estimates are calculated in an alternating order, and wherein a first acoustic path estimate of the plurality of acoustic path estimates and a first clock skew estimate of the plurality of clock skew estimates are used to calculate a second acoustic path estimate of the plurality of acoustic path estimates;
filtering a first portion of the far-end audio with the second acoustic path estimate to generate a replica of echo in the first portion of the far-end audio; and
removing the replica of the echo from a second portion of the near-end audio that corresponds to the first portion of the far-end audio.
2 . The computing system of claim 1 , wherein the plurality of acoustic path estimates are calculated using a Kalman filtering technique.
3 . The computing system of claim 1 , wherein the operations further comprise converting the far-end audio samples and the near-end audio samples from a time domain into a sub-band domain using a multi-hop complex modified discrete cosine transform (MH-CMDCT) with a configurable hop size.
4 . The computing system of claim 3 , wherein the operations further comprise converting sub-domain representations of the far-end audio samples and the near-end audio samples from the sub-band domain to the time domain using an inverse multi-hop complex modified discrete cosine transform (IMH-CMDCT) with the configurable hop size.
5 . A computer-implemented method comprising:
receiving far-end audio samples corresponding to far-end audio that is output from one or more audio output components at a near-end location, wherein the far-end audio is captured at a far-end location and transmitted to the near-end location;
receiving near-end audio samples corresponding to near-end audio that is captured by one or more audio input components at the near-end location;
calculating, using a state-space model, based at least in part on the far-end audio samples and the near-end audio samples, a plurality of acoustic path estimates and a plurality of clock skew estimates, wherein the plurality of acoustic path estimates and the plurality of clock skew estimates are calculated in an alternating order, and wherein a first acoustic path estimate of the plurality of acoustic path estimates and a first clock skew estimate of the plurality of clock skew estimates are used to calculate a second acoustic path estimate of the plurality of acoustic path estimates;
filtering a first portion of the far-end audio with the second acoustic path estimate to generate a replica of echo in the first portion of the far-end audio; and
removing the replica of the echo from a second portion of the near-end audio that corresponds to the first portion of the far-end audio.
6 . The computer-implemented method of claim 5 , wherein the plurality of acoustic path estimates approximate an acoustic path between the one or more audio output components and the one or more audio input components, and wherein the plurality of clock skew estimates approximate a clock skew caused by a difference between a far-end sampling rate associated with the one or more audio output components and a near-end sampling rate associated with the one or more audio input components.
7 . The computer-implemented method of claim 5 , wherein the plurality of acoustic path estimates are calculated using a Kalman filtering technique.
8 . The computer-implemented method of claim 7 , wherein the Kalman filtering technique employs an auxiliary constraint that corresponds to a super-Gaussian distribution.
9 . The computer-implemented method of claim 5 , wherein each acoustic path estimate of the plurality of acoustic path estimates is calculated based at least in part on a preceding acoustic path estimate and a preceding clock skew estimate.
10 . The computer-implemented method of claim 5 , wherein each clock skew estimate of the plurality of clock skew estimates is calculated based at least in part on two preceding acoustic path estimates.
11 . The computer-implemented method of claim 5 , further comprising converting the far-end audio samples and the near-end audio samples from a time domain into a sub-band domain using a multi-hop complex modified discrete cosine transform (MH-CMDCT) with a configurable hop size.
12 . The computer-implemented method of claim 11 , further comprising converting sub-domain representations of the far-end audio samples and the near-end audio samples from the sub-band domain to the time domain using an inverse multi-hop complex modified discrete cosine transform (IMH-CMDCT) with the configurable hop size.
13 . The computer-implemented method of claim 12 , further comprising performing a window tightening procedure to generate tightened windows corresponding to the far-end audio samples and the near-end audio samples for conversion into the sub-band domain and back into the time domain.
14 . The computer-implemented method of claim 12 , wherein different quantities of taps for different sub-bands are used for the calculating of the plurality of acoustic path estimates and the plurality of clock skew estimates.
15 . One or more non-transitory computer-readable storage media having stored thereon computing instructions that, upon execution by one or more computing devices, cause the one or more computing devices to perform computing operations comprising:
receiving far-end audio samples corresponding to far-end audio that is output from one or more audio output components at a near-end location, wherein the far-end audio is captured at a far-end location and transmitted to the near-end location;
receiving near-end audio samples corresponding to near-end audio that is captured by one or more audio input components at the near-end location;
calculating, using a state-space model, based at least in part on the far-end audio samples and the near-end audio samples, a plurality of acoustic path estimates and a plurality of clock skew estimates, wherein the plurality of acoustic path estimates and the plurality of clock skew estimates are calculated in an alternating order, and wherein a first acoustic path estimate of the plurality of acoustic path estimates and a first clock skew estimate of the plurality of clock skew estimates are used to calculate a second acoustic path estimate of the plurality of acoustic path estimates;
filtering a first portion of the far-end audio with the second acoustic path estimate to generate a replica of echo in the first portion of the far-end audio; and
removing the replica of the echo from a second portion of the near-end audio that corresponds to the first portion of the far-end audio.
16 . The one or more non-transitory computer-readable storage media of claim 15 , wherein the plurality of acoustic path estimates are calculated using a Kalman filtering technique.
17 . The one or more non-transitory computer-readable storage media of claim 15 , wherein each acoustic path estimate of the plurality of acoustic path estimates is calculated based at least in part on a preceding acoustic path estimate and a preceding clock skew estimate.
18 . The one or more non-transitory computer-readable storage media of claim 15 , wherein each clock skew estimate of the plurality of clock skew estimates is calculated based at least in part on two preceding acoustic path estimates.
19 . The one or more non-transitory computer-readable storage media of claim 15 , wherein the operations further comprise converting the far-end audio samples and the near-end audio samples from a time domain into a sub-band domain using a multi-hop complex modified discrete cosine transform (MH-CMDCT) with a configurable hop size.
20 . The one or more non-transitory computer-readable storage media of claim 19 , wherein the operations further comprise converting sub-domain representations of the far-end audio samples and the near-end audio samples from the sub-band domain to the time domain using an inverse multi-hop complex modified discrete cosine transform (IMH-CMDCT) with the configurable hop size.