IP Library Granted Patent US 10,354,658
Granted Patent B2
US 10,354,658 · App. 16/173,797 · Granted Jul 16, 2019

Voice control of playback device using voice assistant service(s)

Inventor: Dayn Wilberding (Santa Barbara, CA)
Assignee: Sonos, Inc.
G10L17/22G10L15/22G10L15/30G10L17/02H05B37/02G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,354,658
App. No.
16/173,797
Granted
Jul 16, 2019
Kind
B2
Abstract

Disclosed herein are example techniques to identify a voice service to process a voice input. An example implementation may involve a playback device capturing, via a microphone array, audio into one or more buffers. The playback device analyzes analyzing the captured audio using multiple wake-word detection algorithms. When a particular wake-word detection algorithm detects a wake-word corresponding to a particular voice assistant service, the playback device transmits the captured audio to the particular voice assistant service. The captured audio includes a voice input that includes a command to modify at least one playback setting of a media playback system. After transmitting the captured audio, the playback device receives, from the particular voice assistant service, instructions to modify the at least one playback setting according to the command, modifies the at least one playback setting, and with the at least one playback setting modified, plays back at least one audio track.

Claims (58)

1. A playback device comprising:

one or more amplifiers configured to drive one or more speakers;

a microphone array;

a network interface;

one or more processors;

tangible, non-transitory computer-readable media having stored therein instructions executable by the one or more processors to cause the playback device to perform a method comprising:

continuously capturing, via the microphone array, audio into one or more buffers;

analyzing the captured audio using multiple wake-word detection algorithms running concurrently on the one or more processors, each wake-word detection algorithm corresponding to a respective voice assistant service among multiple voice assistant services supported by the playback device;

when a particular wake-word detection algorithm of the multiple wake-word detection algorithms detects, in the captured audio, a wake-word corresponding to a particular voice assistant service, transmitting, via the network interface, the captured audio to the particular voice assistant service, wherein the captured audio comprises a voice input, wherein the voice input comprises a command to modify at least one playback setting of a media playback system, and wherein the media playback system comprises the playback device;

after transmitting the captured audio, receiving, from one or more servers of the particular voice assistant service via the network interface, instructions to modify the at least one playback setting according to the command;

modifying the at least one playback setting based on the instructions; and

with the at least one playback setting modified, playing back at least one audio track via the one or more amplifiers configured to drive the one or more speakers.

2. The playback device of claim 1 , wherein the method further comprises:

transmitting, via the network interface, a search query, to the particular voice assistant service; and

receiving, from one or more servers of the particular voice assistant service via the network interface in response to the search query, data representing search results, the search results including audio tracks corresponding to the search query, wherein the search results are unique to the particular voice assistant service among the multiple voice assistant services, and wherein the search results comprise the at least one audio track.

3. The playback device of claim 2 , wherein the captured audio is captured first audio, and wherein the method further comprises before capturing the first audio:

continuously capturing, via the microphone array, second audio into the one or more buffers;

analyzing the captured second audio using the multiple wake-word detection algorithms running concurrently on the one or more processors; and

detecting, in the captured second audio, the wake-word corresponding to the particular voice assistant service, wherein the captured second audio comprises a voice command, and wherein the voice command comprises the search query.

4. The playback device of claim 1 , wherein the playback device is a first playback device, wherein modifying the at least one playback setting based on the instructions comprises joining a synchrony group comprising the first playback device and a second playback device, and wherein the method further comprises receiving the at least one audio track from the second playback device via the network interface.

5. The playback device of claim 1 , wherein the playback device is a first playback device, wherein modifying the at least one playback setting based on the instructions comprises forming a synchrony group comprising the first playback device and a second playback device, and wherein playing back the at least one audio track comprises playing back the at least one audio track in synchrony with the second playback device of the synchrony group.

6. The playback device of claim 5 , wherein the method further comprises transmitting the at least one audio track to the second playback device via the network interface.

7. The playback device of claim 1 , wherein modifying the at least one playback setting based on the instructions comprises selecting a music source of the at least one audio track.

8. A tangible, non-transitory computer-readable medium having stored therein instructions executable by one or more processors to cause a playback device to perform a method comprising:

continuously capturing, via a microphone array of the playback device, audio into one or more buffers;

analyzing the captured audio using multiple wake-word detection algorithms running concurrently on one or more processors of the playback device, each wake-word detection algorithm corresponding to a respective voice assistant service among multiple voice assistant services supported by the playback device;

when a particular wake-word detection algorithm of the multiple wake-word detection algorithms detects, in the captured audio, a wake-word corresponding to a particular voice assistant service, transmitting, via a network interface of the playback device, the captured audio to the particular voice assistant service, wherein the captured audio comprises a voice input, wherein the voice input comprises a command to modify at least one playback setting of a media playback system, and wherein the media playback system comprises the playback device;

after transmitting the captured audio, receiving, from one or more servers of the particular voice assistant service via the network interface, instructions to modify the at least one playback setting according to the command;

modifying the at least one playback setting based on the instructions; and

with the at least one playback setting modified, playing back at least one audio track via one or more amplifiers configured to drive one or more speakers.

9. The tangible, non-transitory computer-readable medium of claim 8 , wherein the method further comprises:

transmitting, via the network interface, a search query, to the particular voice assistant service; and

receiving, from one or more servers of the particular voice assistant service via the network interface in response to the search query, data representing search results, the search results including audio tracks corresponding to the search query, wherein the search results are unique to the particular voice assistant service among the multiple voice assistant services, and wherein the search results comprise the at least one audio track.

10. The tangible, non-transitory computer-readable medium of claim 9 , wherein the captured audio is captured first audio, and wherein the method further comprises before capturing the first audio:

continuously capturing, via the microphone array, second audio into the one or more buffers;

analyzing the captured second audio using the multiple wake-word detection algorithms running concurrently on the one or more processors; and

detecting, in the captured second audio, the wake-word corresponding to the particular voice assistant service, wherein the captured second audio comprises a voice command, and wherein the voice command comprises the search query.

11. The tangible, non-transitory computer-readable medium of claim 8 , wherein the playback device is a first playback device, wherein modifying the at least one playback setting based on the instructions comprises joining a synchrony group comprising the first playback device and a second playback device, and wherein the method further comprises receiving the at least one audio track from the second playback device via the network interface.

12. The tangible, non-transitory computer-readable medium of claim 8 , wherein the playback device is a first playback device, wherein modifying the at least one playback setting based on the instructions comprises forming a synchrony group comprising the first playback device and a second playback device, and wherein playing back the at least one audio track comprises playing back the at least one audio track in synchrony with the second playback device of the synchrony group.

13. The tangible, non-transitory computer-readable medium of claim 12 , wherein the method further comprises transmitting the at least one audio track to the second playback device via the network interface.

14. The tangible, non-transitory computer-readable medium of claim 8 , wherein modifying the at least one playback setting based on the instructions comprises selecting a music source of the at least one audio track.

15. A method comprising:

continuously capturing, via a microphone array of a playback device, audio into one or more buffers;

analyzing the captured audio using multiple wake-word detection algorithms running concurrently on one or more processors of the playback device, each wake-word detection algorithm corresponding to a respective voice assistant service among multiple voice assistant services supported by the playback device;

when a particular wake-word detection algorithm of the multiple wake-word detection algorithms detects, in the captured audio, a wake-word corresponding to a particular voice assistant service, transmitting, via a network interface of the playback device, the captured audio to the particular voice assistant service, wherein the captured audio comprises a voice input, wherein the voice input comprises a command to modify at least one playback setting of a media playback system, and wherein the media playback system comprises the playback device;

after transmitting the captured audio, receiving, from one or more servers of the particular voice assistant service via the network interface, instructions to modify the at least one playback setting according to the command;

modifying the at least one playback setting based on the instructions; and

with the at least one playback setting modified, playing back at least one audio track via one or more amplifiers configured to drive one or more speakers.

16. The method of claim 15 , further comprising:

transmitting, via the network interface, a search query, to the particular voice assistant service; and

receiving, from one or more servers of the particular voice assistant service via the network interface in response to the search query, data representing search results, the search results including audio tracks corresponding to the search query, wherein the search results are unique to the particular voice assistant service among the multiple voice assistant services, and wherein the search results comprise the at least one audio track.

17. The method of claim 16 , wherein the captured audio is captured first audio, and wherein the method further comprises before capturing the first audio:

continuously capturing, via the microphone array, second audio into the one or more buffers;

analyzing the captured second audio using the multiple wake-word detection algorithms running concurrently on the one or more processors; and

detecting, in the captured second audio, the wake-word corresponding to the particular voice assistant service, wherein the captured second audio comprises a voice command, and wherein the voice command comprises the search query.

18. The method of claim 15 , wherein the playback device is a first playback device, wherein modifying the at least one playback setting based on the instructions comprises joining a synchrony group comprising the first playback device and a second playback device, and wherein the method further comprises receiving the at least one audio track from the second playback device via the network interface.

19. The method of claim 15 , wherein the playback device is a first playback device, wherein modifying the at least one playback setting based on the instructions comprises forming a synchrony group comprising the first playback device and a second playback device, and wherein playing back the at least one audio track comprises playing back the at least one audio track in synchrony with the second playback device of the synchrony group.

20. The method of claim 15 , wherein modifying the at least one playback setting based on the instructions comprises selecting a music source of the at least one audio track.

Assignments (2)
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 29, 2018
From: WILBERDING, DAYN
To: SONOS, INC.
Reel/Frame 047343/0441 →
Continuity (2)
Continuation 15229868 · Aug 5, 2016
Related Publication 20190074014A1 · Mar 7, 2019
Cited By (33)
US 12,211,490 US 12,212,945 US 12,217,748 US 12,217,765 US 12,230,291 US 12,236,932 US 12,250,536 US 12,265,746 US 12,277,368 US 12,279,096 US 12,283,269 US 12,288,558 US 12,314,633 US 12,322,390 US 12,327,549 US 12,327,556 US 12,360,734 US 12,375,052 US 12,387,716 US 12,424,220 US 12,438,977 US 12,462,802 US 12,482,467 US 12,505,832 US 12,513,479 US 12,518,755 US 12,518,756 US 12,579,978 US 12,614,549 US 12,626,717 US 12,640,148 US 12,699,543 US 12,711,962