IP Library Granted Patent US 10,482,868
Granted Patent B2
US 10,482,868 · App. 15/718,911 · Granted Nov 19, 2019

Multi-channel acoustic echo cancellation

Inventors: Saeed Bagheri Sereshki (Goleta, CA); Romi Kadri (Santa Barbara, CA)
Assignee: Sonos, Inc.
G10K11/178G06F3/165G10L21/0208H04B17/336H04L65/601H04M9/082H04R27/00G10K2210/3012G10K2210/505G10L2021/02082H04R2227/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,482,868
App. No.
15/718,911
Granted
Nov 19, 2019
Kind
B2
Abstract

A method of operating a playback device includes receiving source audio content that includes a first and second channel stream of audio. The method also includes playing back, via a first and second speaker driver of the playback device, the first and second channel streams of audio, thereby producing a first and second channel audio output. A captured stream of audio is received by a microphone of the playback device, and portions of the captured stream of audio correspond to the first and second channel audio outputs. The first and second channel streams of audio are combined into a compound audio signal, and acoustic echo cancellation is performed on the compound audio signal to produce an acoustic echo cancellation output, which is then applied to the captured stream of audio to increase the signal-to noise ratio of the captured stream of audio.

Claims (53)

1. A method of operating a playback device having a first speaker driver, at least a second speaker driver, and one or more microphones, the method comprising:

receiving, via a network interface of the playback device, a source stream of audio comprising source audio content to be played back by the playback device, wherein the source audio content comprises a first channel stream of audio and a second channel stream of audio;

producing a first channel audio output by playing back, via the first speaker driver, the first channel stream of audio;

producing a second channel audio output by playing back, via the second speaker driver, the second channel stream of audio;

receiving, via the one or more microphones, a captured stream of audio comprising a first portion corresponding to the first channel audio output and a second portion corresponding to the second channel audio output, wherein the captured stream of audio has a first signal-to-noise ratio;

performing a singular value decomposition on the first channel stream of audio and the second channel stream of audio to result in a combined set of signal components;

selecting a subset of the combined set of signal components based on one or more parameters;

performing acoustic echo cancellation on the subset of the combined set of signal components, wherein performing acoustic echo cancellation on the subset of the combined set of signal components produces a first acoustic echo cancellation output; and

applying the first acoustic echo cancellation output to the captured stream of audio, thereby increasing the signal-to noise ratio of the captured stream of audio from the first signal-to-noise ratio to a second signal-to-noise ratio, wherein the second signal-to-noise ratio is greater than the first signal-to-noise ratio.

2. The method of claim 1 , wherein the captured stream of audio comprises a third portion corresponding to a vocal command issued by a user, and wherein applying the first acoustic echo cancellation output to the captured stream of audio, thereby increasing the signal-to-noise ratio of the captured stream of audio from the first signal-to-noise ratio to the second signal-to-noise ratio, results in the first portion and second portion being eliminated or minimized in the captured stream of audio.

3. The method of claim 1 , further comprising:

transforming both of the subset of the combined set of signal components and the captured stream of audio into a Short-Time Fourier Transform domain.

4. The method of claim 1 , wherein selecting the subset of the combined set of signal components based on the one or more parameters comprises selecting a subset of the combined set of signal components having at least one of (a) an energy content above a first threshold energy content or (b) a calculated variance above a first threshold variance.

5. The method of claim 1 ,

wherein the playback device comprises a third speaker driver,

wherein the source audio content comprises the first channel stream of audio, the second channel stream of audio, and a third channel stream of audio, and

wherein the method further comprises producing a third channel audio output by playing back, via the third speaker driver, the third channel stream of audio.

6. The method of claim 5 , wherein performing the singular value decomposition on the first channel stream of audio and the second channel stream of audio to result in the combined set of signal components comprises:

performing the singular value decomposition on the first channel stream of audio, the second channel stream of audio, and the third channel stream of audio to result in the combined set of signal components.

7. The method of claim 1 , further comprising:

detecting a trigger to perform acoustic echo cancellation on the subset of the combined set of signal components, wherein detecting the trigger comprises detecting that (a) a playback function is initiated by the playback device or (b) an unmute command is received by the playback device after the playback function is initiated.

8. The method of claim 1 , wherein performing the singular value decomposition on the first channel stream of audio is performed by one or more processors of the playback device,

wherein performing the singular value decomposition on the second channel stream of audio is performed by the one or more processors of the playback device,

wherein selecting the subset of the combined set of signal components based on the one or more parameters is performed by the one or more processors of the playback device,

wherein performing acoustic echo cancellation on the subset of the combined set of signal components is performed by the one or more processors of the playback device, and

wherein applying the first acoustic echo cancellation output to the captured stream of audio is performed by the one or more processors of the playback device.

9. A tangible, non-transitory computer-readable medium storing instructions that when executed by one or more processors cause a playback device to perform functions comprising:

receiving, via a network interface of the playback device, a source stream of audio comprising source audio content to be played back by the playback device, wherein the playback device comprises a first speaker driver and at least a second speaker driver and further comprises one or more microphones, and wherein the source audio content comprises a first channel stream of audio and a second channel stream of audio;

producing a first channel audio output by playing back, via the first speaker driver, the first channel stream of audio;

producing a second channel audio output by playing back, via the second speaker driver, the second channel stream of audio;

receiving, via the one or more microphones, a captured stream of audio comprising a first portion corresponding to the first channel audio output, and further comprising a second portion corresponding to the second channel audio output, wherein the captured stream of audio has a first signal-to-noise ratio;

performing a singular value decomposition on the first channel stream of audio and the second channel stream of audio to result in a combined set of signal components;

selecting a subset of the combined set of signal components based on one or more parameters;

performing acoustic echo cancellation on the subset of the combined set of signal components, wherein performing acoustic echo cancellation on the subset of the combined set of signal components produces a first acoustic echo cancellation output; and

applying the first acoustic echo cancellation output to the captured stream of audio, thereby increasing the signal-to noise ratio of the captured stream of audio from the first signal-to-noise ratio to a second signal-to-noise ratio, wherein the second signal-to-noise ratio is greater than the first signal-to-noise ratio.

10. The computer-readable medium of claim 9 , wherein the captured stream of audio comprises a third portion corresponding to a vocal command issued by a user, and wherein applying the first acoustic echo cancellation output to the captured stream of audio, thereby increasing the signal-to-noise ratio of the captured stream of audio from the first signal-to-noise ratio to the second signal-to-noise ratio, results in the first portion and second portion being eliminated or minimized in the captured stream of audio.

11. The computer-readable medium of claim 9 , further comprising instructions that when executed by the one or more processors cause the playback device to perform functions comprising:

transforming both of the subset of the combined set of signal components and the captured stream of audio into a Short-Time Fourier Transform domain.

12. The computer-readable medium of claim 9 , wherein selecting the subset of the combined set of signal components based on the one or more parameters comprises selecting a subset of the combined set of signal components having at least one of (a) an energy content above a first threshold energy content or (b) a calculated variance above a first threshold variance.

13. The computer-readable medium of claim 9 ,

wherein the playback device comprises a third speaker driver,

wherein the source audio content comprises the first channel stream of audio, the second channel stream of audio, and a third channel stream of audio, and

wherein the method further comprises producing a third channel audio output by playing back, via the third speaker driver, the third channel stream of audio.

14. The computer-readable medium of claim 13 , wherein performing the singular value decomposition on the first channel stream of audio and the second channel stream of audio to result in the combined set of signal components comprises:

performing the singular value decomposition on the first channel stream of audio, the second channel stream of audio, and the third channel stream of audio to result in the combined set of signal components.

15. The computer-readable medium of claim 9 , further comprising instructions that when executed by the one or more processors cause the playback device to perform functions comprising:

detecting a trigger to perform acoustic echo cancellation on the subset of the first combined set of signal components, wherein detecting the trigger comprises detecting that (a) a playback function is initiated by the playback device or (b) an unmute command is received by the playback device after the playback function is initiated.

16. The computer-readable medium of claim 9 ,

wherein performing the singular value decomposition on the first channel stream of audio is performed by the one or more processors,

wherein performing the singular value decomposition on the second channel stream of audio is performed by the one or more processors,

wherein selecting the subset of the combined set of signal components based on the one or more parameters is performed by the one or more processors,

wherein performing acoustic echo cancellation on the subset of the combined set of signal components is performed by the one or more processors, and

wherein applying the first acoustic echo cancellation output to the captured stream of audio is performed by the one or more processors.

Assignments (4)
RELEASE OF SECURITY INTEREST Recorded Oct 18, 2021
From: JPMORGAN CHASE BANK, N.A.
To: SONOS, INC.
Reel/Frame 058213/0597 →
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
SECURITY INTEREST Recorded Aug 30, 2018
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 046991/0433 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2017
From: SERESHKI, SAEED BAGHERI; KADRI, ROMI
To: SONOS, INC.
Reel/Frame 044399/0265 →
Continuity (1)
Related Publication 20190096384A1 · Mar 28, 2019