IP Library Granted Patent US 12705020
Granted Patent B2
US 12705020 · App. 19/109,920 · Granted Aug 11, 2026

Primary-ambient playback on audio playback devices

Inventor: Christopher Pike (Manchester, GB)
Assignee: Sonos, Inc.
G06F3/165H04R1/02H04R5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12705020
App. No.
19/109,920
Granted
Aug 11, 2026
Kind
B2
Abstract

Example techniques relate to primary-ambient playback of surround audio by audio playback devices. Example playback devices described herein may include multiple speakers, such as a forward-firing audio transducer and side-firing audio transducers. Using example techniques, such playback devices may perform primary-ambient decomposition to separate surround channels into primary′ and ambient channels. The playback devices play back the ambient channel(s) via the side-firing transducers and the primary channel via the forward-firing transducer.

Claims (73)

1 . A system comprising:

a first playback device including a forward-firing audio transducer and side-firing audio transducers;

a second playback device including an additional forward-firing audio transducer and additional side-firing audio transducers;

a third playback device;

a network interface;

at least one processor; and

data storage including instructions that are executable by the at least one processor such that the system is configured to perform functions comprising:

receiving data representing multi-channel audio content, the multi-channel audio content comprising a first audio signal representing a first surround channel, a second audio signal representing a second surround channel, and one or more third audio signals;

performing primary-ambient decomposition on the first audio signal and the second audio signal to generate (i) a first primary signal and a first ambient signal and (ii) a second primary signal and a second ambient signal, wherein performing the primary-ambient decomposition on the first audio signal and the second audio signal comprises:

determining first signals representing the first audio signal in respective frequency bands in the short time Fourier transform (STFT) domain, wherein each first signal comprises a respective series of first segments representing the first audio signal in respective time periods;

determining second signals representing the second audio signal in respective frequency bands in the STFT domain, wherein each second signal comprises a respective series of second segments representing the second audio signal in the respective time periods; and

performing the primary-ambient decomposition on the first signals and the second signals;

causing the first playback device to play back the first primary signal via first forward-firing audio transducer and the first ambient signal via the side-firing audio transducers in synchrony with playback of the one or more third audio signals by the third playback device; and

causing the second playback device to play back the second primary signal via the additional forward-firing audio transducer and the second ambient signal via the additional side-firing audio transducers in synchrony with playback of the one or more third audio signals by the third playback device.

2 . The system of claim 1 , wherein the side-firing audio transducers of the first playback device comprise a first side-firing audio transducer and a second side-firing audio transducer, and wherein the first playback device comprises a housing carrying the first side-firing audio transducer, the second side-firing audio transducer, and the forward-firing audio transducer such that the first side-firing audio transducer is directed in a first direction, the second side-firing audio transducer is directed in a second direction that is opposite the first direction, and the forward-firing audio transducer is directed in a third direction that is orthogonal to the first direction and the second direction.

3 . The system of claim 2 , wherein causing the first playback device to play back the first primary signal via the forward-firing audio transducer comprises:

playing back a portion of the first primary signal via the forward-firing audio transducer and at least one of the side-firing audio transducers such that a beam is formed that having a primary lobe in a direction between the first direction and the second direction.

4 . The system of claim 2 , wherein the functions further comprise:

detecting an obstruction within a minimum proximity to the first side-firing audio transducer; and

adjusting a mix between the first side-firing audio transducer and the second side-firing audio transducer to output more of the first ambient signal via the first side-firing audio transducer.

5 . The system of claim 2 , wherein the functions further comprise:

detecting obstructions within a minimum proximity to the first side-firing audio transducer and the second side-firing audio transducer; and

adjusting a mix between the side-firing audio transducers and the forward-firing audio transducer to output at least a portion of the first ambient signal via the forward-firing audio transducer.

6 . The system of claim 1 , wherein the multi-channel audio content includes an object-based mix comprising objects represented by metadata, and wherein causing the first playback device to play back the first primary signal via the forward-firing audio transducer comprises:

playing back a portion of the first primary signal via the forward-firing audio transducer and at least one of the side-firing audio transducers to represent one or more objects according to the metadata.

7 . The system of claim 1 , wherein the primary-ambient decomposition adapts decomposition of the first signals and the seconds signals over multiple first segments and second segments, and wherein performing the primary-ambient decomposition on the first signals and the second signals comprises:

during playback of the first ambient signal, analyzing the first segments with a transient detector; and

when the transient detector detects a transient in the first segments, adapting decomposition of the first signals to decompose one or more particular first segments including the transient to the first primary signal.

8 . The system of claim 1 , wherein the third playback device comprises a high-definition multimedia interface, and wherein the third playback device comprises the at least one processor and the data storage, and wherein receiving the multi-channel audio content comprises receiving the multi-channel audio content via the high-definition multimedia interface, and wherein the functions further comprise:

causing the third playback device to play back the one or more third audio signals.

9 . The system of claim 1 , wherein the first playback device comprises the network interface, the at least one processor and the data storage, and wherein receiving the multi-channel audio content comprises receiving the multi-channel audio content via the network interface.

10 . The system of claim 1 , wherein the functions further comprise:

buffering the received data representing the multi-channel audio content into one or more buffers, and wherein performing the primary-ambient decomposition comprises:

as the received multi-channel audio content is buffered, performing the primary-ambient decomposition.

11 . A first playback device comprising:

a forward-firing audio transducer and side-firing audio transducers;

a network interface;

at least one processor; and

data storage including instructions that are executable by the at least one processor such that the first playback device is configured to perform functions comprising,

receiving data representing multi-channel audio content, the multi-channel audio content comprising a first audio signal representing a first surround channel a second audio signal representing a second surround channel;

performing primary-ambient decomposition on the first audio signal and the second audio signal to generate primary signal and a first ambient signal and (ii) a second primary signal and a second ambient signal, wherein performing the primary-ambient decomposition on the first audio signal and the second audio signal comprises:

determining first signals representing the first audio signal in respective frequency bands in the short time Fourier transform (STFT) domain, wherein each first signal comprises a respective series of first segments representing the first audio signal in respective time periods;

determining second signals representing the second audio signal in respective frequency bands in the STFT domain, wherein each second signal comprises a respective series of second segments representing the second audio signal in the respective time periods; and

performing the primary-ambient decomposition on the first signals and the second signals;

causing a second playback device to play back the second primary signal via an additional forward-firing audio transducer and the second ambient signal via an additional side-firing audio transducers in synchrony with playback of one or more third audio signals by a third playback device; and

playing back the first primary signal via the forward-firing audio transducer and the first ambient signal via the side-firing audio transducers in synchrony with playback of the one or more third audio signals by the third playback device.

12 . The first playback device of claim 11 , wherein the side-firing audio transducers of the first playback device comprise a first side-firing audio transducer and a second side-firing audio transducer, and wherein the first playback device comprises a housing carrying the first side-firing audio transducer, the second side-firing audio transducer, and the forward-firing audio transducer such that the first side-firing audio transducer is directed in a first direction, the second side-firing audio transducer is directed in a second direction that is opposite the first direction, and the forward-firing audio transducer is directed in a third direction that is orthogonal to the first direction and the second direction.

13 . The first playback device of claim 12 , wherein playing back the first primary signal via the forward-firing audio transducer comprises:

playing back a portion of the first primary signal via the forward-firing audio transducer and at least one of the side-firing audio transducers such that a beam is formed that having a primary lobe in a direction between the first direction and the second direction.

14 . The first playback device of claim 12 , wherein the functions further comprise:

detecting an obstruction within a minimum proximity to the first side-firing audio transducer; and

adjusting a mix between the first side-firing audio transducer and the second side-firing audio transducer to output more of the first ambient signal via the first side-firing audio transducer.

15 . The first playback device of claim 12 , wherein the functions further comprise:

detecting obstructions within a minimum proximity to the first side-firing audio transducer and the second side-firing audio transducer; and

adjusting a mix between the side-firing audio transducers and the forward-firing audio transducer to output at least a portion of the first ambient signal via the forward-firing audio transducer.

16 . The first playback device of claim 11 , wherein the multi-channel audio content includes an object-based mix comprising objects represented by metadata, and wherein causing the first playback device to play back the first primary signal via the forward-firing audio transducer comprises:

playing back a portion of the first primary signal via the forward-firing audio transducer and at least one of the side-firing audio transducers to represent one or more objects according to the metadata.

17 . The first playback device of claim 11 , wherein the primary-ambient decomposition adapts decomposition of the first signals and the seconds signals over multiple first segments and second segments, and wherein performing the primary-ambient decomposition on the first signals and the second signals comprises:

during playback of the first ambient signal, analyzing the first segments with a transient detector; and

when the transient detector detects a transient in the first segments, adapting decomposition of the first signals to decompose one or more particular first segments including the transient to the first primary signal.

18 . The first playback device of claim 11 , wherein the first playback device comprises the network interface, the at least one processor and the data storage, and wherein receiving the multi-channel audio content comprises receiving the multi-channel audio content via the network interface.

19 . The first playback device of claim 11 , wherein the functions further comprise:

buffering the received data representing the multi-channel audio content into one or more buffers, and wherein performing the primary-ambient decomposition comprises:

as the received multi-channel audio content is buffered, performing the primary-ambient decomposition.

20 . A method comprising:

receiving data representing multi-channel audio content, the multi-channel audio content comprising a first audio signal representing a first surround channel, a second audio signal representing a second surround channel, and one or more third audio signals;

buffering the received data representing the multi-channel audio content into one or more buffers;

as the received multi-channel audio content is buffered, performing primary-ambient decomposition on the first audio signal and the second audio signal to generate (i) a first primary signal and a first ambient signal and (ii) a second primary signal and a second ambient signal, wherein performing the primary-ambient decomposition on the first audio signal and the second audio signal comprises:

determining first signals representing the first audio signal in respective frequency bands in the short time Fourier transform (STFT) domain, wherein each first signal comprises a respective series of first segments representing the first audio signal in respective time periods;

determining second signals representing the second audio signal in respective frequency bands in the STFT domain, wherein each second signal comprises a respective series of second segments representing the second audio signal in the respective time periods; and

performing the primary-ambient decomposition on the first signals and the second signals:

causing a first playback device to play back the first primary signal via a first forward-firing audio transducer and the first ambient signal via-side-firing audio transducers in synchrony with playback of the one or more third audio signals by a second playback device; and

causing a third playback device to play back the second primary signal via an additional forward-firing audio transducer and the second ambient signal via additional side-firing audio transducers in synchrony with playback of the one or more third audio signals by the second playback device.