IP Library Granted Patent US 10,565,998
Granted Patent B2
US 10,565,998 · App. 16/437,437 · Granted Feb 18, 2020

Playback device supporting concurrent voice assistant services

Inventor: Dayn Wilberding (Santa Barbara, CA)
Assignee: Sonos, Inc.
G10L17/22G06F3/167G10L15/22G10L15/30G10L17/02H05B37/02G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,565,998
App. No.
16/437,437
Granted
Feb 18, 2020
Kind
B2
Abstract

Disclosed herein are example techniques to support multiple voice assistant services. An example implementation may involve a playback device continuously capturing, via the at least one microphone, audio into one or more buffers and analyzing the captured audio using a first wake-word detection algorithm and a second wake-word detection algorithm. When one of the first wake-word detection algorithm or the second wake-word detection algorithm detects, in the captured audio, a wake-word corresponding to a particular voice assistant service of (a) the first voice assistant service or (b) the second voice assistant service, the playback device transmits the captured audio to one or more servers associated with the particular voice assistant service. After transmitting the captured audio, the playback device receives, via the network interface, at least one instruction based on the captured audio; and performs one or more actions based on the at least one instruction.

Claims (109)

1. A playback device comprising:

one or more amplifiers configured to drive one or more speakers;

at least one microphone;

a network interface;

one or more processors; and

data storage having stored therein instructions executable by the one or more processors to cause the playback device to perform a method comprising:

registering the playback device with a first voice assistant service;

after registering the playback device with the first voice assistant service, receiving from a computing device, an instruction to register the playback device with a second voice assistant service;

after receiving the instruction to register the playback device with the second voice assistant service, registering the playback device with the second voice assistant service such that the playback device is concurrently registered to the first and second voice assistant services;

continuously capturing, via the at least one microphone, audio into one or more buffers;

analyzing the captured audio using a first wake-word detection algorithm and a second wake-word detection algorithm, wherein the first wake-word detection algorithm corresponds to a first wake word associated with the first voice assistant service, and wherein the second wake-word detection algorithm corresponds to a second wake word associated with the second voice assistant service;

when one of the first wake-word detection algorithm and the second wake-word detection algorithm detects, in the captured audio, a wake word corresponding to a particular voice assistant service of (a) the first voice assistant service or (b) the second voice assistant service, transmitting the captured audio to one or more servers associated with the particular voice assistant service;

after transmitting the captured audio, receiving, via the network interface, at least one instruction based on the captured audio; and

after receiving the at least one instruction, performing one or more actions based on the at least one instruction.

2. The playback device of claim 1 , wherein the method further comprises:

assigning the first voice assistant service as a default voice assistant service;

further capturing, via the at least one microphone, audio into the one or more buffers;

analyzing the further captured audio using the first wake-word detection algorithm and the second wake-word detection algorithm;

detecting, in the further captured audio, the second wake word;

determining that the second voice assistance service is unavailable to process the further captured audio;

after determining that the second voice assistance service is unavailable to process the further captured audio, transmitting the further captured audio to one or more servers associated with the default voice assistant service;

after transmitting the further captured audio, receiving from the default voice assistant service, via the network interface, at least one instruction based on the further captured audio; and

after receiving the at least one instruction from the default voice assistant service, performing one or more actions based on the at least one instruction from the default voice assistant service.

3. The playback device of claim 1 , wherein receiving from the computing device an instruction to register the playback device with the second voice assistant service comprises receiving the instruction from a remote computing device associated with the second voice assistant service.

4. The playback device of claim 1 , wherein performing the one or more actions comprises receiving audio via the network interface and playing back the received audio via the one or more amplifiers configured to drive the one or more speakers.

5. The playback device of claim 4 , wherein the particular voice assistant service comprises the first voice assistant service, and wherein the method further comprises:

further capturing, via the at least one microphone, audio into the one or more buffers;

analyzing the further captured audio using the first wake-word detection algorithm and the second wake-word detection algorithm;

detecting, via the second wake-word detection algorithm, the second wake word in the further captured audio data;

after detecting the second wake word, transmitting the further captured audio to one or more servers associated with the second voice assistant service, wherein the further captured audio comprises a query;

receiving, via the network interface in response to the query, data corresponding to results of the query; and

playing back audio based on the data via the one or more amplifiers configured to drive the one or more speakers.

6. The playback device of claim 1 , wherein performing the one or more actions comprise modifying at least one playback setting of a media playback system, wherein the media playback system comprises the playback device.

7. The playback device of claim 6 , wherein the particular voice assistant service comprises the first voice assistant service, and wherein the method further comprises:

further capturing, via the at least one microphone, audio into the one or more buffers;

analyzing the further captured audio using the first wake-word detection algorithm and the second wake-word detection algorithm;

detecting, via the second wake-word detection algorithm, the second wake word in the further captured audio data;

after detecting the second wake word, transmitting the further captured audio to one or more servers associated with the second voice assistant service, wherein the further captured audio comprises a voice command to play particular audio;

receiving, via the network interface in response to the voice command, instructions to play back at least one audio track; and

after receiving the instructions to play back the at least one audio track, playing back the at least audio track via the one or more amplifiers configured to drive the one or more speakers.

8. A method to be performed by a playback device comprising a network interface, at least one microphone, and one or more amplifiers configured to drive one or more speakers, the method comprising:

registering the playback device with a first voice assistant service;

after registering the playback device with the first voice assistant service, receiving from a computing device an instruction to register the playback device with a second voice assistant service;

after receiving the instruction to register the playback device with the second voice assistant service, registering the playback device with the second voice assistant service such that the playback device is concurrently registered to the first and second voice assistant services;

continuously capturing, via the at least one microphone, audio into one or more buffers;

analyzing the captured audio using a first wake-word detection algorithm and a second wake-word detection algorithm, wherein the first wake-word detection algorithm corresponds to a first wake word associated with the first voice assistant service, and wherein the second wake-word detection algorithm corresponds to a second wake word associated with the second voice assistant service;

when one of the first wake-word detection algorithm and the second wake-word detection algorithm detects, in the captured audio, a wake word corresponding to a particular voice assistant service of (a) the first voice assistant service or (b) the second voice assistant service, transmitting the captured audio to one or more servers associated with the particular voice assistant service;

after transmitting the captured audio, receiving, via the network interface, at least one instruction based on the captured audio; and

after receiving the at least one instruction, performing one or more actions based on the at least one instruction.

9. The method of claim 8 , further comprising:

assigning the first voice assistant service as a default voice assistant service;

further capturing, via the at least one microphone, audio into the one or more buffers;

analyzing the further captured audio using the first wake-word detection algorithm and the second wake-word detection algorithm;

detecting, in the further captured audio, the second wake word;

determining that the second voice assistance service is unavailable to process the further captured audio;

after determining that the second voice assistance service is unavailable to process the further captured audio, transmitting the further captured audio to one or more servers associated with the default voice assistant service;

after transmitting the further captured audio, receiving, from the one or more servers of the default voice assistant service via the network interface, at least one instruction based on the further captured audio; and

after receiving the at least one instruction from the default voice assistant service, performing one or more actions based on the at least one instruction from the default voice assistant service.

10. The method of claim 8 , wherein receiving from the computing device an instruction to register the playback device with the second voice assistant service comprises receiving the instruction from a remote computing device associated with the second voice assistant service.

11. The method of claim 8 , wherein performing the one or more actions comprises receiving audio via the network interface and playing back the received audio via the one or more amplifiers configured to drive the one or more speakers.

12. The method of claim 11 , wherein the particular voice assistant service comprises the first voice assistant service, and wherein the method further comprises:

further capturing, via the at least one microphone, audio into the one or more buffers;

analyzing the further captured audio using the first wake-word detection algorithm and the second wake-word detection algorithm;

detecting, via the second wake-word detection algorithm, the second wake word in the further captured audio data;

after detecting the second wake word, transmitting the further captured audio to one or more servers associated with the second voice assistant service, wherein the further captured audio comprises a query;

receiving, via the network interface in response to the query, data corresponding to results of the query; and

playing back audio based on the data via the one or more amplifiers configured to drive the one or more speakers.

13. The method of claim 8 , wherein performing the one or more actions comprise modifying at least one playback setting of a media playback system, wherein the media playback system comprises the playback device.

14. The method of claim 13 , wherein the particular voice assistant service comprises the first voice assistant service, and wherein the method further comprises:

further capturing, via the at least one microphone, audio into the one or more buffers;

analyzing the further captured audio using the first wake-word detection algorithm and the second wake-word detection algorithm;

detecting, via the second wake-word detection algorithm, the second wake word in the further captured audio data;

after detecting the second wake word, transmitting the further captured audio to one or more servers associated with the second voice assistant service, wherein the further captured audio comprises a voice command to play particular audio;

receiving, via the network interface in response to the voice command, instructions to play back at least one audio track; and

after receiving the instructions to play back the at least one audio track, playing back the at least audio track via the one or more amplifiers configured to drive the one or more speakers.

15. A non-transitory computer-readable medium having instructions stored thereon that are executable by one or more processors to cause a playback device to perform a method, the playback device comprising a network interface, at least one microphone, and one or more amplifiers configured to drive one or more speakers, the method comprising:

registering the playback device with a first voice assistant service;

after registering the playback device with the first voice assistant service, receiving from a computing device an instruction to register the playback device with a second voice assistant service;

after receiving the instruction to register the playback device with the second voice assistant service, registering the playback device with the second voice assistant service such that the playback device is concurrently registered to the first and second voice assistant services;

continuously capturing, via the at least one microphone, audio into one or more buffers;

analyzing the captured audio using a first wake-word detection algorithm and a second wake-word detection algorithm, wherein the first wake-word detection algorithm corresponds to a first wake word associated with the first voice assistant service, and wherein the second wake-word detection algorithm corresponds to a second wake word associated with the second voice assistant service;

when one of the first wake-word detection algorithm and the second wake-word detection algorithm detects, in the captured audio, a wake word corresponding to a particular voice assistant service of (a) the first voice assistant service or (b) the second voice assistant service, transmitting the captured audio to one or more servers associated with the particular voice assistant service;

after transmitting the captured audio, receiving, via the network interface, at least one instruction based on the captured audio; and

after receiving the at least one instruction, performing one or more actions based on the at least one instruction.

16. The non-transitory computer-readable medium of claim 15 , wherein the method further comprises:

assigning the first voice assistant service as a default voice assistant service;

further capturing, via the at least one microphone, audio into the one or more buffers;

analyzing the further captured audio using the first wake-word detection algorithm and the second wake-word detection algorithm;

detecting, in the further captured audio, the second wake word;

determining that the second voice assistance service is unavailable to process the further captured audio;

after determining that the second voice assistance service is unavailable to process the further captured audio, transmitting the further captured audio to one or more servers associated with the default voice assistant service;

after transmitting the further captured audio, receiving from the default voice assistant service, via the network interface, at least one instruction based on the further captured audio; and

after receiving the at least one instruction from the default voice assistant service, performing one or more actions based on the at least one instruction from the default voice assistant service.

17. The non-transitory computer-readable medium of claim 15 , wherein receiving from the computing device an instruction to register the playback device with the second voice assistant service comprises receiving the instruction from a remote computing device associated with the second voice assistant service.

18. The non-transitory computer-readable medium of claim 15 , wherein performing the one or more actions comprises receiving audio via the network interface and playing back the received audio via the one or more amplifiers configured to drive the one or more speakers.

19. The non-transitory computer-readable medium of claim 18 , wherein the particular voice assistant service comprises the first voice assistant service, and wherein the method further comprises:

further capturing, via the at least one microphone, audio into the one or more buffers;

analyzing the further captured audio using the first wake-word detection algorithm and the second wake-word detection algorithm;

detecting, via the second wake-word detection algorithm, the second wake word in the further captured audio data;

after detecting the second wake word, transmitting the further captured audio to one or more servers associated with the second voice assistant service, wherein the further captured audio comprises a query;

receiving, via the network interface in response to the query, data corresponding to results of the query; and

playing back audio based on the data via the one or more amplifiers configured to drive the one or more speakers.

20. The non-transitory computer-readable medium of claim 15 , wherein performing the one or more actions comprise modifying at least one playback setting of a media playback system, wherein the media playback system comprises the playback device, wherein the particular voice assistant service comprises the first voice assistant service, and wherein the method further comprises:

further capturing, via the at least one microphone, audio into the one or more buffers;

analyzing the further captured audio using the first wake-word detection algorithm and the second wake-word detection algorithm;

detecting, via the second wake-word detection algorithm, the second wake word in the further captured audio data;

after detecting the second wake word, transmitting the further captured audio to one or more servers associated with the second voice assistant service, wherein the further captured audio comprises a voice command to play particular audio;

receiving, via the network interface in response to the voice command, instructions to play back at least one audio track; and

after receiving the instructions to play back the at least one audio track, playing back the at least audio track via the one or more amplifiers configured to drive the one or more speakers.

Assignments (2)
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2019
From: WILBERDING, DAYN
To: SONOS, INC.
Reel/Frame 049432/0384 →
Continuity (3)
Continuation 16173797 · Oct 29, 2018
Continuation 15229868 · Aug 5, 2016
Related Publication 20190295555A1 · Sep 26, 2019
Cited By (32)
US 12,211,490 US 12,212,945 US 12,217,748 US 12,217,765 US 12,230,291 US 12,236,932 US 12,265,746 US 12,277,368 US 12,279,096 US 12,283,269 US 12,288,558 US 12,314,633 US 12,322,390 US 12,327,549 US 12,327,556 US 12,360,734 US 12,375,052 US 12,387,716 US 12,424,220 US 12,438,977 US 12,462,802 US 12,482,467 US 12,505,832 US 12,513,479 US 12,518,755 US 12,518,756 US 12,579,978 US 12,614,549 US 12,626,717 US 12,640,148 US 12,699,543 US 12,711,962