IP Library › Granted Patent US 11,817,076
Granted Patent B2
US 11,817,076 · App. 18/145,501 · Granted Nov 14, 2023

Multi-channel acoustic echo cancellation

Inventors: Saeed Bagheri Sereshki (Goleta, CA); Romi Kadri (Santa Barbara, CA)
Assignee: Sonos, Inc.
G10K11/178G06F3/165G10L21/0208H04B17/336H04L65/75H04M9/082H04R27/00G10K2210/3012G10K2210/505G10L2021/02082H04R2227/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,817,076
App. No.
18/145,501
Granted
Nov 14, 2023
Kind
B2
Abstract

A playback device is configured to: produce a first channel audio output of a first channel of audio content; produce a second channel audio output of a second channel of the audio content; receive captured audio content comprising (i) a first portion corresponding to the first channel audio output, (ii) a second portion corresponding to the second channel audio output, and (iii) a third portion corresponding to a voice command, wherein the captured audio content has a first signal-to-noise ratio; determine a set of signal components from at least one of the first channel or the second channel of the audio content; perform acoustic echo cancellation on a subset of signal components; determine an acoustic echo cancellation output; and apply the acoustic echo cancellation output to the captured audio content and thereby increase the first signal-to-noise ratio to a second signal-to-noise ratio that is greater than the first signal-to-noise ratio.

Claims (64)

1. A playback device comprising:

a first set of one or more transducers;

a second set of one of more transducers;

at least one processor;

a network interface;

a non-transitory computer-readable medium; and

program instructions stored on the non-transitory computer-readable medium that are executable by the at least one processor such that the playback device is configured to:

produce, via the first set of one or more transducers, a first channel audio output of a first channel of given audio content;

produce, via the second set of one or more transducers, a second channel audio output of a second channel of the given audio content;

receive, by one or more microphones, captured audio content comprising (i) a first portion corresponding to the first channel audio output, (ii) a second portion corresponding to the second channel audio output, and (iii) a third portion corresponding to a voice command, wherein the captured audio content has a first signal-to-noise ratio;

determine a set of signal components from at least one of the first channel or the second channel of the given audio content;

select a subset of signal components from the set of signal components;

perform acoustic echo cancellation on the subset of signal components and thereby determine an acoustic echo cancellation output; and

apply the acoustic echo cancellation output to the captured audio content and thereby increase a signal-to-noise ratio of the captured audio content from the first signal-to-noise ratio to a second signal-to-noise ratio that is greater than the first signal-to-noise ratio.

2. The playback device of claim 1 , further comprising program instructions stored on the non-transitory computer-readable medium that are executable by the at least one processor such that the playback device is configured to:

before producing the first and second channel audio outputs, receive audio content for playback by the playback device, wherein the received audio content comprises the first channel of the given audio content and the second channel of the given audio content.

3. The playback device of claim 1 , wherein the playback device comprises the one or more microphones.

4. The playback device of claim 1 , wherein the program instructions that are executable by the at least one processor such that the playback device is configured to determine the set of signal components comprise program instructions that are executable by the at least one processor such that the playback device is configured to:

determine the set of signal components based on a singular value decomposition of the first and second channels of the given audio content.

5. The playback device of claim 4 , wherein each of the subset of signal components has at least one of (i) an energy content above a threshold energy content or (ii) a calculated variance above a threshold variance.

6. The playback device of claim 1 , wherein the program instructions that are executable by the at least one processor such that the playback device is configured to determine the set of signal components comprise program instructions that are executable by the at least one processor such that the playback device is configured to:

determine the set of signal components based on a cross-correlation of the first and second channels of the given audio content.

7. The playback device of claim 1 , wherein the second signal-to-noise ratio is greater than the first signal-to-noise ratio by a range of 1 dB to 10 db.

8. The playback device of claim 1 , further comprising:

a third set of one or more transducers;

a fourth set of one or more transducers; and

program instructions stored on the non-transitory computer-readable medium that are executable by the at least one processor such that the playback device is configured to:

produce a third channel audio output by playing back, via the third set of one or more transducers, a third channel of the given audio content; and

produce a fourth channel audio output by playing back, via the fourth set of one or more transducers, a fourth channel of the given audio content.

9. The playback device of claim 1 , wherein application of the acoustic echo cancellation output to the captured audio content results in the first portion and the second portion being eliminated or minimized.

10. The playback device of claim 1 , wherein:

the playback device further comprises program instructions stored on the non-transitory computer-readable medium that are executable by the at least one processor such that the playback device is configured to:

combine at least the first channel of the given audio content and the second channel of the given audio content into a compound audio signal; and

the program instructions that are executable by the at least one processor such that the playback device is configured to determine the set of signal components comprise program instructions that are executable by the at least one processor such that the playback device is configured to:

determine the set of signal components from the compound audio signal.

11. The playback device of claim 1 , further comprising program instructions stored on the non-transitory computer-readable medium that are executable by the at least one processor such that the playback device is configured to:

before performing the acoustic echo cancellation on the subset of signal components, transform the captured audio content and at least one of the first channel of the given audio content or the second channel of the given audio content from a time domain into a frequency domain.

12. A non-transitory computer-readable medium, wherein the non-transitory computer-readable medium is provisioned with program instructions that, when executed by at least one processor, cause a playback device to:

produce, via a first set of one or more transducers, a first channel audio output of a first channel of given audio content;

produce, via a second set of one or more transducers, a second channel audio output of a second channel of the given audio content;

receive, by one or more microphones, captured audio content comprising (i) a first portion corresponding to the first channel audio output, (ii) a second portion corresponding to the second channel audio output, and (iii) a third portion corresponding to a voice command, wherein the captured audio content has a first signal-to-noise ratio;

determine a set of signal components from at least one of the first channel or the second channel of the given audio content;

select a subset of signal components from the set of signal components;

perform acoustic echo cancellation on the subset of signal components and thereby determine an acoustic echo cancellation output; and

apply the acoustic echo cancellation output to the captured audio content and thereby increase a signal-to-noise ratio of the captured audio content from the first signal-to-noise ratio to a second signal-to-noise ratio that is greater than the first signal-to-noise ratio.

13. The non-transitory computer-readable medium of claim 12 , wherein the non-transitory computer-readable medium is also provisioned with program instructions that, when executed by at least one processor, cause the playback device to:

before producing the first and second channel audio outputs, receive audio content for playback by the playback device, wherein the received audio content comprises the first channel of the given audio content and the second channel of the given audio content.

14. The non-transitory computer-readable medium of claim 12 , wherein the playback device comprises the one or more microphones.

15. The non-transitory computer-readable medium of claim 12 , wherein the program instructions that, when executed by at least one processor, cause the playback device to determine the set of signal components comprise program instructions that, when executed by at least one processor, cause the playback device to:

determine the set of signal components based on a singular value decomposition of the first and second channels of the given audio content.

16. The non-transitory computer-readable medium of claim 15 , wherein each of the subset of signal components has at least one of (i) an energy content above a threshold energy content or (ii) a calculated variance above a threshold variance.

17. The non-transitory computer-readable medium of claim 12 , wherein the program instructions that, when executed by at least one processor, cause the playback device to determine the set of signal components comprise program instructions that, when executed by at least one processor, cause the playback device to:

determine the set of signal components based on a cross-correlation of the first and second channels of the given audio content.

18. A method carried out by a playback device, the method comprising:

producing, via a first set of one or more transducers, a first channel audio output of a first channel of given audio content;

producing, via a second set of one or more transducers, a second channel audio output of a second channel of the given audio content;

receiving, by one or more microphones, captured audio content comprising (i) a first portion corresponding to the first channel audio output, (ii) a second portion corresponding to the second channel audio output, and (iii) a third portion corresponding to a voice command, wherein the captured audio content has a first signal-to-noise ratio;

determining a set of signal components from at least one of the first channel or the second channel of the given audio content;

selecting a subset of signal components from the set of signal components;

performing acoustic echo cancellation on the subset of signal components and thereby determining an acoustic echo cancellation output; and

applying the acoustic echo cancellation output to the captured audio content and thereby increasing a signal-to-noise ratio of the captured audio content from the first signal-to-noise ratio to a second signal-to-noise ratio that is greater than the first signal-to-noise ratio.

19. The method of claim 18 , further comprising:

before producing the first and second channel audio outputs, receiving audio content for playback by the playback device, wherein the received audio content comprises the first channel of the given audio content and the second channel of the given audio content.

20. The method of claim 18 , wherein the playback device comprises the one or more microphones.

Assignments (2)
SECURITY INTEREST Recorded Jan 30, 2026
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074533/0615 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2023
From: SERESHKI, SAEED BAGHERI; KADRI, ROMI
To: SONOS, INC.
Reel/Frame 063796/0059 →
Continuity (4)
Continuation 17145667 · Jan 11, 2021
Continuation 16598125 · Oct 10, 2019
Continuation 15718911 · Sep 28, 2017
Related Publication 20230127040A1 · Apr 27, 2023
Cited By (34)
US 12,192,713 US 12,210,801 US 12,211,490 US 12,217,748 US 12,217,765 US 12,230,291 US 12,231,859 US 12,236,932 US 12,288,558 US 12,314,633 US 12,322,390 US 12,340,802 US 12,360,734 US 12,374,334 US 12,375,052 US 12,424,220 US 12,438,977 US 12,462,802 US 12,498,899 US 12,505,832 US 12,513,466 US 12,513,479 US 12,518,755 US 12,518,756 US 12,562,167 US 12,578,779 US 12,579,978 US 12,626,717 US 12,699,543 US 12,711,962 US 12,730,606 US 12,732,547 US 12,744,035 US 12,749,486