IP Library Granted Patent US 11,736,860
Granted Patent B2
US 11,736,860 · App. 17/562,412 · Granted Aug 22, 2023

Voice control of a media playback system

Inventors: Jonathan P. Lang (Santa Barbara, CA); Mark Plagge (Santa Barbara, CA); Simon Jarvis (Santa Barbara, CA); Romi Kadri (Cambridge, MA); Yean-Nian Willy Chen (Santa Barbara, CA); Paul Andrew Bates (Santa Barbara, CA); Luis Vega-Zayas (Cambridge, MA); Christopher Butts (Evanston, IL); Nicholas A. J. Millington (Santa Barbara, CA); Keith Corbin (Santa Barbara, CA)
Assignee: Sonos, Inc.
H04R3/00G06F3/162G06F3/165G06F3/167G10L15/14G10L15/22H04L12/2803H04L12/2809H04R3/12H04R27/00H04R29/007H04S7/301H04S7/303H04W8/005H04W8/24G10L21/02G10L2015/223H04L2012/2849H04R2227/003H04R2227/005H04R2420/07H04W84/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,736,860
App. No.
17/562,412
Granted
Aug 22, 2023
Kind
B2
Abstract

Multiple aspects of systems and methods for voice control and related features and functionality for various embodiments of media playback devices, networked microphone devices, microphone-equipped media playback devices, and speaker-equipped networked microphone devices are disclosed and described herein, including but not limited to designating and managing default networked devices, audio response playback, room-corrected voice detection, content mixing, music service selection, metadata exchange between networked playback systems and networked microphone systems, handling loss of pairing between networked devices, actions based on user identification, and other voice control of networked devices.

Claims (76)

1. A system comprising:

at least one processor;

at least one tangible, non-transitory computer-readable medium; and

program instructions stored on the at least one tangible, non-transitory computer-readable medium that are executable by the at least one processor such that the system is configured to:

obtain metadata from a network computing device relating to a configuration of a media playback system, wherein the metadata indicates that (i) a first playback device is configured to operate in a first playback zone and (ii) the first playback device and a second playback device are configured to operate in a second playback zone;

cause the first playback device to operate in the first playback zone in a given playback state comprising play back of one or more media items identified in a playback queue associated with the first playback zone;

while the first playback device is operating in the given playback state:

receive data corresponding to a detected voice input, wherein the data comprises an indication within the voice input of (i) a command word and (ii) one or more zone variable instances; and

determine, based on the command word and the one or more zone variable instances, an intent to transfer the given playback state to the second playback zone; and

after determining the intent to transfer the given playback state to the second playback zone, transfer the given playback state to the second playback zone, thereby causing the second playback device in the second playback zone to play back the one or more media items identified in the playback queue.

2. The system of claim 1 , wherein:

the voice input does not identify any playback device of the media playback system that is to execute a command corresponding to the voice input.

3. The system of claim 1 , further comprising:

at least one microphone, wherein the program instructions are executable by the at least one processor such that the system is configured to detect the voice input via the at least one microphone.

4. The system of claim 3 , wherein one of the first playback device and the second playback device comprise the at least one microphone.

5. The system of claim 3 , wherein the program instructions are executable by the at least one processor such that the system is configured to:

receive an indication of a direction of the voice input received via the at least one microphone; and

direct audio output of at least one of the first playback device and the second playback device based on the indication of the direction of the voice input.

6. The system of claim 1 , wherein the program instructions are executable by the at least one processor such that the system is configured to:

play back the one or more media items at a first volume level;

play back audio content associated with a response to the voice input at a second volume level; and

adjust playback of the one or more media items at a third volume level for at least a duration of the playback of the audio content associated with the response to the voice input, wherein the third volume level is lower than each of the first volume level and the second volume level.

7. The system of claim 1 , wherein at least a portion of the playback queue is stored on a remote computing device associated with a cloud-based computing system.

8. A system comprising:

a first playback device configured to communicate over at least one data network, wherein the first playback device comprises:

at least one first processor;

at least one first tangible, non-transitory computer-readable medium; and

first program instructions stored on the at least one first tangible, non-transitory computer-readable medium that are executable by the at least one first processor such that the first playback device is configured to operate in a given playback state comprising play back of one or more media items identified in a playback queue associated with a first playback zone; and

at least one computing device configured to communicate over the at least one data network, wherein the at least one computing device comprises:

at least one second processor;

at least one second tangible non-transitory computer-readable medium;

second program instructions stored on the at least one second tangible, non-transitory computer-readable medium that are executable by the at least one second processor of the at least one computing device such that the at least one computing device is configured to:

obtain metadata relating to a configuration of a media playback system, wherein the metadata indicates that (i) the first playback device is configured to operate in the first playback zone and (ii) the first playback device and a second playback device are configured to operate in a second playback zone;

while the first playback device is operating in the given playback state:

receive data corresponding to a detected voice input, wherein the data comprises an indication within the voice input of (i) a command word and (ii) one or more zone variable instances; and

determine, based on the command word and the one or more zone variable instances, an intent to transfer the given playback state to the second playback zone; and

after determining the intent to transfer the given playback state to the second playback zone, transfer the given playback state to the second playback zone, thereby causing the second playback device in the second playback zone to play back the one or more media items identified in the playback queue.

9. The system of claim 8 , wherein

the voice input does not identify any playback device of the media playback system that is to execute a command corresponding to the voice input.

10. The system of claim 8 , further comprising:

at least one microphone, wherein the second program instructions are executable by the at least one second processor of the at least one computing device such that the system is configured to detect the voice input via the at least one microphone.

11. The system of claim 10 , wherein one of the first playback device and the second playback device comprise the at least one microphone.

12. The system of claim 10 , wherein the first program instructions are executable by the at least one first processor of the first playback device such that the system is configured to:

receive an indication of a direction of the voice input received via the at least one microphone; and

direct audio output of at least one of the first playback device and the second playback device based on the indication of the direction of the voice input.

13. The system of claim 10 , wherein the first program instructions are executable by the at least one first processor of the first playback device such that the system is configured to:

play back the one or more media items at a first volume level;

play back audio content associated with a response to the voice input at a second volume level; and

adjust playback of the one or more media items at a third volume level for at least a duration of the playback of the audio content associated with the response to the voice input, wherein the third volume level is lower than each of the first volume level and the second volume level.

14. A system comprising a first playback device and a second playback device each configured to communicate over at least one data network,

wherein the first playback device comprises:

at least one first processor;

at least one first tangible, non-transitory computer-readable medium; and

first program instructions stored on the at least one first tangible, non-transitory computer-readable medium that are executable by the at least one first processor such that the first playback device is configured to:

obtain metadata from a network computing device relating to a configuration of a media playback system, wherein the metadata indicates that (i) the first playback device is configured to operate in a first playback zone and (ii) the first playback device and the second playback device are configured to operate in a second playback zone;

operate in a given playback state comprising play back of one or more media items identified in a playback queue associated with the first playback zone; and

while the first playback device is operating in the given playback state:

receive data corresponding to a detected voice input, wherein the data comprises an indication within the voice input of (i) a command word and (ii) one or more zone variable instances;

determine, based on the command word and the one or more zone variable instances, an intent to transfer the given playback state to the second playback zone; and

after determining the intent to transfer the given playback state to the second playback zone, transfer the given playback state to the second playback zone; and

wherein the second playback device comprises:

at least one second processor;

at least one second tangible, non-transitory computer-readable medium; and

second program instructions stored on the at least one second tangible, non-transitory computer-readable medium of the second playback device that are executable by the at least one second processor such that the second playback device is configured to play back the one or more media items identified in the playback queue after transferring the given playback state to the second playback zone.

15. The system of claim 14 , wherein

the voice input does not identify any playback device of the media playback system that is to execute a command corresponding to the voice input.

16. The system of claim 14 , wherein the first playback device comprises at least one first microphone, wherein the first program instructions are executable by the at least one first processor of the first playback device such that the system is configured to detect the voice input via the at least one first microphone.

17. The system of claim 14 , wherein the second playback device comprises at least one second microphone, wherein the second program instructions are executable by the at least one second processor of the second playback device such that the system is configured to detect the voice input via the at least one second microphone.

18. The system of claim 14 , wherein the first program instructions are executable by the at least one first processor of the first playback device such that the system is configured to:

receive an indication of a direction of the voice input received via at least one microphone; and

direct audio output of at least one of the first playback device and the second playback device based on the indication of the direction of the voice input.

19. The system of claim 14 , wherein the first program instructions are executable by the at least one first processor of the first playback device such that the system is configured to:

play back the one or more media items at a first volume level;

play back audio content associated with a response to the voice input at a second volume level; and

adjust playback of the one or more media items at a third volume level for at least a duration of the playback of the audio content associated with the response to the voice input, wherein the third volume level is lower than each of the first volume level and the second volume level.

20. The system of claim 14 , wherein at least a portion of the playback queue is stored on a remote computing device associated with a cloud-based computing system.

Assignments (3)
SECURITY INTEREST Recorded Jan 30, 2026
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074533/0615 →
CORRECTIVE ASSIGNMENT TO CORRECT THE APPLICATION NUMBER FROM 17564412 TO 17562412 PREVIOUSLY RECORDED ON REEL 63479 FRAME 28. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Dec 12, 2025
From: LANG, JONATHAN P.; PLAGGE, MARK; JARVIS, SIMON; KADRI, ROMI; CHEN, YEAN-NIAN WILLY; BATES, PAUL ANDREW; VEGA-ZAYAS, LUIS; BUTTS, CHRISTOPHER; MILLINGTON, NICHOLAS A. J.; CORBIN, KEITH
To: SONOS, INC.
Reel/Frame 073952/0774 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2023
From: LANG, JONATHAN P.; PLAGGE, MARK; JARVIS, SIMON; KADRI, ROMI; CHEN, YEAN-NIAN WILLY; BATES, PAUL ANDREW; VEGA-ZAYAS, LUIS; BUTTS, CHRISTOPHER; MILLINGTON, NICHOLAS A.J.; CORBIN, KEITH
To: SONOS, INC.
Reel/Frame 063479/0777 →
Continuity (12)
Continuation 17008104 · Aug 31, 2020
Continuation 16700607 · Dec 2, 2019
Continuation 15438749 · Feb 21, 2017
Provisional Application 62298410 · Feb 22, 2016
Provisional Application 62298418 · Feb 22, 2016
Provisional Application 62298433 · Feb 22, 2016
Provisional Application 62298439 · Feb 22, 2016
Provisional Application 62298425 · Feb 22, 2016
Provisional Application 62298350 · Feb 22, 2016
Provisional Application 62298388 · Feb 22, 2016
Provisional Application 62298393 · Feb 22, 2016
Related Publication 20230054164A1 · Feb 23, 2023
Cited By (2)
US 12,573,407 US 12,602,197