IP Library Granted Patent US 12,634,397
Granted Patent B1
US 12,634,397 · App. 18/061,738 · Granted May 19, 2026

Clock skew robust acoustic echo cancellation

Inventors: Karim Helwani (Mountain View, CA); Erfan Soltanmohammadi (Silver Spring, MD); Michael Mark Goodwin (Scotts Valley, CA); Arvindh Krishnaswamy (Palo Alto, CA)
Assignee: Amazon Technologies, Inc.
H04M9/082H04M9/087
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,634,397
App. No.
18/061,738
Granted
May 19, 2026
Kind
B1
Abstract

Far-end audio samples may be received corresponding to far-end audio that is output from one or more audio output components. Near-end audio samples may be received corresponding to near-end audio that is captured by one or more audio input components. A plurality of acoustic path estimates and a plurality of clock skew estimates may be calculated in an alternating order, using a state-space model, based at least in part on the far-end audio samples and the near-end audio samples. A first acoustic path estimate and a first clock skew estimate may be used to calculate a second acoustic path estimate. A first portion of the far-end audio may be filtered with the second acoustic path estimate to generate a replica of echo in the first portion of the far-end audio. The replica of the echo may be removed from a corresponding second portion of the near-end audio.

Claims (37)

1 . A computing system comprising:

one or more processors; and

one or more memories having stored therein instructions that, upon execution by the one or more processors, cause the computing system to perform computing operations comprising:

receiving far-end audio samples corresponding to far-end audio that is output from one or more audio output components at a near-end location, wherein the far-end audio is captured at a far-end location and transmitted to the near-end location;

receiving near-end audio samples corresponding to near-end audio that is captured by one or more audio input components at the near-end location;

calculating, using a state-space model, based at least in part on the far-end audio samples and the near-end audio samples, a plurality of acoustic path estimates and a plurality of clock skew estimates, wherein the plurality of acoustic path estimates approximate an acoustic path between the one or more audio output components and the one or more audio input components, and wherein the plurality of clock skew estimates approximate a clock skew caused by a difference between a far-end sampling rate associated with the one or more audio output components and a near-end sampling rate associated with the one or more audio input components, wherein the plurality of acoustic path estimates and the plurality of clock skew estimates are calculated in an alternating order, and wherein a first acoustic path estimate of the plurality of acoustic path estimates and a first clock skew estimate of the plurality of clock skew estimates are used to calculate a second acoustic path estimate of the plurality of acoustic path estimates;

filtering a first portion of the far-end audio with the second acoustic path estimate to generate a replica of echo in the first portion of the far-end audio; and

removing the replica of the echo from a second portion of the near-end audio that corresponds to the first portion of the far-end audio.

2 . The computing system of claim 1 , wherein the plurality of acoustic path estimates are calculated using a Kalman filtering technique.

3 . The computing system of claim 1 , wherein the operations further comprise converting the far-end audio samples and the near-end audio samples from a time domain into a sub-band domain using a multi-hop complex modified discrete cosine transform (MH-CMDCT) with a configurable hop size.

4 . The computing system of claim 3 , wherein the operations further comprise converting sub-domain representations of the far-end audio samples and the near-end audio samples from the sub-band domain to the time domain using an inverse multi-hop complex modified discrete cosine transform (IMH-CMDCT) with the configurable hop size.

5 . A computer-implemented method comprising:

receiving far-end audio samples corresponding to far-end audio that is output from one or more audio output components at a near-end location, wherein the far-end audio is captured at a far-end location and transmitted to the near-end location;

receiving near-end audio samples corresponding to near-end audio that is captured by one or more audio input components at the near-end location;

calculating, using a state-space model, based at least in part on the far-end audio samples and the near-end audio samples, a plurality of acoustic path estimates and a plurality of clock skew estimates, wherein the plurality of acoustic path estimates and the plurality of clock skew estimates are calculated in an alternating order, and wherein a first acoustic path estimate of the plurality of acoustic path estimates and a first clock skew estimate of the plurality of clock skew estimates are used to calculate a second acoustic path estimate of the plurality of acoustic path estimates;

filtering a first portion of the far-end audio with the second acoustic path estimate to generate a replica of echo in the first portion of the far-end audio; and

removing the replica of the echo from a second portion of the near-end audio that corresponds to the first portion of the far-end audio.

6 . The computer-implemented method of claim 5 , wherein the plurality of acoustic path estimates approximate an acoustic path between the one or more audio output components and the one or more audio input components, and wherein the plurality of clock skew estimates approximate a clock skew caused by a difference between a far-end sampling rate associated with the one or more audio output components and a near-end sampling rate associated with the one or more audio input components.

7 . The computer-implemented method of claim 5 , wherein the plurality of acoustic path estimates are calculated using a Kalman filtering technique.

8 . The computer-implemented method of claim 7 , wherein the Kalman filtering technique employs an auxiliary constraint that corresponds to a super-Gaussian distribution.

9 . The computer-implemented method of claim 5 , wherein each acoustic path estimate of the plurality of acoustic path estimates is calculated based at least in part on a preceding acoustic path estimate and a preceding clock skew estimate.

10 . The computer-implemented method of claim 5 , wherein each clock skew estimate of the plurality of clock skew estimates is calculated based at least in part on two preceding acoustic path estimates.

11 . The computer-implemented method of claim 5 , further comprising converting the far-end audio samples and the near-end audio samples from a time domain into a sub-band domain using a multi-hop complex modified discrete cosine transform (MH-CMDCT) with a configurable hop size.

12 . The computer-implemented method of claim 11 , further comprising converting sub-domain representations of the far-end audio samples and the near-end audio samples from the sub-band domain to the time domain using an inverse multi-hop complex modified discrete cosine transform (IMH-CMDCT) with the configurable hop size.

13 . The computer-implemented method of claim 12 , further comprising performing a window tightening procedure to generate tightened windows corresponding to the far-end audio samples and the near-end audio samples for conversion into the sub-band domain and back into the time domain.

14 . The computer-implemented method of claim 12 , wherein different quantities of taps for different sub-bands are used for the calculating of the plurality of acoustic path estimates and the plurality of clock skew estimates.

15 . One or more non-transitory computer-readable storage media having stored thereon computing instructions that, upon execution by one or more computing devices, cause the one or more computing devices to perform computing operations comprising:

receiving far-end audio samples corresponding to far-end audio that is output from one or more audio output components at a near-end location, wherein the far-end audio is captured at a far-end location and transmitted to the near-end location;

receiving near-end audio samples corresponding to near-end audio that is captured by one or more audio input components at the near-end location;

calculating, using a state-space model, based at least in part on the far-end audio samples and the near-end audio samples, a plurality of acoustic path estimates and a plurality of clock skew estimates, wherein the plurality of acoustic path estimates and the plurality of clock skew estimates are calculated in an alternating order, and wherein a first acoustic path estimate of the plurality of acoustic path estimates and a first clock skew estimate of the plurality of clock skew estimates are used to calculate a second acoustic path estimate of the plurality of acoustic path estimates;

filtering a first portion of the far-end audio with the second acoustic path estimate to generate a replica of echo in the first portion of the far-end audio; and

removing the replica of the echo from a second portion of the near-end audio that corresponds to the first portion of the far-end audio.

16 . The one or more non-transitory computer-readable storage media of claim 15 , wherein the plurality of acoustic path estimates are calculated using a Kalman filtering technique.

17 . The one or more non-transitory computer-readable storage media of claim 15 , wherein each acoustic path estimate of the plurality of acoustic path estimates is calculated based at least in part on a preceding acoustic path estimate and a preceding clock skew estimate.

18 . The one or more non-transitory computer-readable storage media of claim 15 , wherein each clock skew estimate of the plurality of clock skew estimates is calculated based at least in part on two preceding acoustic path estimates.

19 . The one or more non-transitory computer-readable storage media of claim 15 , wherein the operations further comprise converting the far-end audio samples and the near-end audio samples from a time domain into a sub-band domain using a multi-hop complex modified discrete cosine transform (MH-CMDCT) with a configurable hop size.

20 . The one or more non-transitory computer-readable storage media of claim 19 , wherein the operations further comprise converting sub-domain representations of the far-end audio samples and the near-end audio samples from the sub-band domain to the time domain using an inverse multi-hop complex modified discrete cosine transform (IMH-CMDCT) with the configurable hop size.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2024
From: KRISHNASWAMY, ARVINDH
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 066823/0377 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2022
From: SOLTANMOHAMMADI, ERFAN; GOODWIN, MICHAEL MARK; HELWANI, KARIM
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 061979/0589 →
References Cited (45)
US 4677668A · Ardalan · 1987 [cited by examiner]
US 8879438B2 · Saleem · 2014 [cited by examiner]
US 9083783B2 · Ikram · 2015 [cited by examiner]
US 9343073B1 · Murgia · 2016 [cited by examiner]
US 9444566B1 · Mustiere · 2016 [cited by examiner]
US 9671822B2 · Aweya · 2017 [cited by examiner]
US 9916840B1 · Do · 2018 [cited by examiner]
US 10374786B1 · Mustiere · 2019 [cited by examiner]
US 11018789B2 · Aweya · 2021 [cited by examiner]
US 11509411B2 · Vincent · 2022 [cited by examiner]
US 11875810B1 · Helwani · 2024 [cited by examiner]
US 20130044873A1 · Etter · 2013 [cited by examiner]
US 20180145863A1 · Chaloupka · 2018 [cited by examiner]
US 20240048931A1 · Southwell · 2024 [cited by examiner]
US 20240056757A1 · Southwell · 2024 [cited by examiner]
US 20240204895A1 · Wang · 2024 [cited by examiner]
J. Benesty; “Adaptive Estimation of Clock Skew and Different Types of Delay in the Internet Network”; Adaptive Signal Processing; 2003; p. 341-351. [cited by applicant]
Wang et al.; “Correlation Maximization-Based Sampling Rate Offset Estimation for Distributed Microphone Arrays”; IEEE/ACM Transactions on Audio, Speech, and Language Processing; vol. 24; Mar. 2016; p. 571-582. [cited by applicant]
Cherkassky et al.; “Blind Synchronization in Wireless Acoustic Sensor Networks”; IEEE/ACM Transactions on Audio, Speech, and Language Processing; vol. 25 No. 3; Mar. 2017; p. 651-661. [cited by applicant]
Thune et al.; “Tracking Theory of Adaptive Filters with Input-Output Sampling Rate Offset”; 27 [cited by applicant]
C. Avendano; “Acoustic echo suppression in the STFT domain”; IEEE Workshop on the Acoustics of Signal Processing to Audio and Acoustics; Oct. 2001; p. 175-178. [cited by applicant]
Helwani et al.; “A Single-Channel MVDR Filter for Acoustic Echo Suppression”; IEEE Signal Processing Letters; vol. 20; Apr. 2013; p. 351-354. [cited by applicant]
Valin et al.; “Low-Complexity, Real-Time Joint Neural Echo Control and Speech Enhancement Based On Percepnet”; IEEE Int'l Conf. on Acoustics, Speech and Signal Processing; 2021; p. 7133-7137. [cited by applicant]
Fa-Long Lou; Mobile Multimedia Broadcasting Standards; Springer; © 2009; 674 pages. [cited by applicant]
Mathew et al.; “Modified MP3 encoder using complex modified cosine transform”; Int'l Conf. on Multimedia and Expo; vol. 2; 2003; p. 709-712. [cited by applicant]
S. Waldron; An Introduction to Finite Tight Frames; Springer; © 2018; 567 pages. [cited by applicant]
Humphreys et al.; “Kalman Filtering with Newton's Method [Lecture Notes]”; IEEE Control Systems Magazine; vol. 30; Dec. 2010; p. 101-106. [cited by applicant]
Buchner et al.; “Blind Signal Processing for Time-varying Convolutive Mixing Systems Based on Sequence Estimation on Partly Smooth Manifolds”; IEEE Int'l Conf. on Acoustics, Speech and Signal Processing; 2019; p. 7913-7… [cited by applicant]
S. Haykin; Adaptive Filter Theory; 4 [cited by applicant]
P. Huber; Robust Statistics; John Wiley & Sons; 2004; 308 pages. [cited by applicant]
Gansler et al.; “Double-talk robust fast converging algorithms for network echo cancellation”; IEEE Workshop on Applications of Signal Processing to Audio and Acoustics; Oct. 1999; p. 215-218. [cited by applicant]
Buchner et al.; “Robust extended multidelay filter and double-talk detector for acoustic echo cancellation”; IEEE Transactions on Audio, Speech and Language Processing; vol. 14; Sep. 2006; p. 1633-1644. [cited by applicant]
Jin et al.; “Algorithms for robust linear regression by exploiting the connection to sparse signal recovery”; IEEE Int'l Conf. on Acoustics, Speech and Signal Processing; 2010; p. 3830-3833. [cited by applicant]
P.862—Perceptual evaluation of speech quality (PESQ): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs; Int'l Telecommunication Union; 2001; 30 pages. [cited by applicant]
Cutler et al.; “ICASSP 2022 Acoustic Echo Cancellation Challenge”; IEEE Int'l Conf. on Acoustics, Speech and Signal Processing; Feb. 2022; 5 pages. [cited by applicant]
Soo et al.; “Multidelay block frequency domain adaptive filter”; IEEE Transactions on Acoustics, Speech and Signal Processing; vol. 38 No. 2; Feb. 1990; p. 373-376. [cited by applicant]
Kuech et al.; “State-space architecture of the partitioned-block-based acoustic echo controller”; IEEE Int'l Conf. on Acoustics, Speech and Signal Processing; 2014; p. 1309-1313. [cited by applicant]
Dietzen et al.; “Partitioned block frequency domain Kalman filter for multi-channel linear prediction based blind speech dereverberation”; IEEE Int'l Workshop on Acoustic Signal Enhancement; 2016; 5 pages. [cited by applicant]
Kellermann et al.; “Acoustic echo cancellation in subbands”; The Journal of the Acoustical Society of America; vol. 87; 1990; p. S2. [cited by applicant]
Valin et al.; “On Adjusting the Learning Rate in Frequency Domain Echo Cancellation With Double-Talk”; IEEE Transactions on Audio, Speech and Language Processing; vol. 15; Mar. 2007; p. 1030-1034. [cited by applicant]
Ochiai et al.; “Echo Canceler with Two Echo Path Models”; IEEE Transactions on Communications; vol. 25 No. 6; Jun. 1977; p. 589-595. [cited by applicant]
Tai et al.; “Audio Watermarking over the Air with Modulated Self-correlation”; IEEE Int'l Conf. on Acoustics, Speech and Signal Processing; 2019; p. 2452-2456. [cited by applicant]
Enzner et al.; “Frequency-domain adaptive Kalman filter for acoustic echo control in hands-free telephones”; Signal Processing; vol. 86; 2006; p. 1140-1156. [cited by applicant]
Buchner et al.; “Unsupervised Bayesian Estimation and Tracking of Time-Varying Convolutive Multichannel Systems”; 22th Int'l Conf. on Information Fusion; 2019; 8 pages. [cited by applicant]
“Wikipedia—Toeplitz matrix,” webpage <https://en.wikipedia.org/wiki/Toeplitz_matrix#Discrete_convolution> saved on Oct. 18, 2022 by Internet Archive Wayback Machine, retrieved from Internet Archive Wayback Machine <http… [cited by applicant]