IP Library Granted Patent US 10,847,164
Granted Patent B2
US 10,847,164 · App. 16/790,621 · Granted Nov 24, 2020

Playback device supporting concurrent voice assistants

Inventor: Dayn Wilberding (Santa Barbara, CA)
Assignee: Sonos, Inc.
G10L17/22G06F3/167G10L15/22G10L15/30G10L17/02H05B47/10G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,847,164
App. No.
16/790,621
Granted
Nov 24, 2020
Kind
B2
Abstract

Disclosed herein are example techniques to support multiple voice assistant services. An example implementation may involve a playback device capturing audio from the one or more microphones into one or more buffers as a sound data stream monitoring the sound data stream for a wake word associated with a specific voice assistant service and monitoring the sound data stream for a wake word associated with the media playback system. The playback device generates a second wake-word event corresponding to a voice input when sound data matching the wake word associated with the media playback system in a portion of the sound data stream is detected. The playback device determines that the voice input includes sound data matching one or more playback commands and sends sound data representing the voice input to a voice assistant associated with the media playback system for processing of the second voice input.

Claims (76)

1. A playback device of a media playback system, the playback device comprising:

one or more microphones;

a network interface;

one or more processors; and

data storage having stored therein instructions executable by the one or more processors to cause the playback device to perform functions comprising:

capturing audio from the one or more microphones into one or more buffers as a sound data stream;

monitoring, via a first wake word engine, the sound data stream from the one or more microphones for a wake word associated with a specific voice assistant service;

monitoring, via a second wake word engine, the sound data stream from the one or more microphones for a wake word associated with the media playback system;

generating a first wake-word event corresponding to a first voice input when the first wake-word engine detects sound data matching the wake word associated with the specific voice assistant service in a first portion of the sound data stream;

based on generating the first wake-word event, sending sound data representing the first voice input to the specific voice assistant for processing of the first voice input;

generating a second wake-word event corresponding to a second voice input when the second wake-word engine detects sound data matching the wake word associated with the media playback system in a second portion of the sound data stream;

determining that the second voice input includes sound data matching one or more playback commands; and

based on (i) generating the second wake-word event and (ii) determining that the second voice input includes sound data matching one or more playback commands, sending sound data representing the second voice input to a voice assistant associated with the media playback system for processing of the second voice input.

2. The playback device of claim 1 , wherein sending the sound data representing the second voice input to the voice assistant associated with the media playback system comprises streaming, via the network interface, the sound data representing the second voice input to one or more remote servers associated with the media playback system.

3. The playback device of claim 1 , wherein the functions further comprise:

receiving, from the voice assistant associated with the media playback system, at least one playback instruction corresponding to the one or more playback commands; and

performing the at least one playback instruction.

4. The playback device of claim 3 , wherein the voice assistant associated with the media playback system identifies specific media content based on the second voice input, and wherein the at least one playback instruction instructs the playback device to play back the specific media content.

5. The playback device of claim 1 , wherein the functions further comprise:

generating a third wake-word event corresponding to a third voice input when the second wake-word engine detects sound data matching the wake word associated with the media playback system in a third portion of the sound data stream;

determining that the third voice input includes sound data matching a particular playback command; and

based on (i) generating the third wake-word event and (ii) determining that the third voice input includes sound data matching the particular playback command, performing the particular playback command.

6. The playback device of claim 1 , wherein the functions further comprise:

generating a third wake-word event corresponding to a third voice input when the second wake-word engine detects sound data matching the wake word associated with the media playback system in a third portion of the sound data stream;

determining that the third voice input includes sound data matching one or more smart home commands; and

based on (i) generating the third wake-word event and (ii) determining that the third voice input includes sound data matching one or more smart home commands, causing one or more smart home devices to carry out at least one smart home instruction corresponding to the one or more smart home commands.

7. The playback device of claim 6 , wherein causing the one or more smart home devices to carry out at least one smart home instruction corresponding to the one or more smart home commands comprises:

sending sound data representing the third voice input to a voice assistant for processing of the third voice input.

8. A method to be performed by a playback device of a media playback system, the method comprising:

capturing audio from one or more microphones into one or more buffers as a sound data stream;

monitoring, via a first wake word engine, the sound data stream from the one or more microphones for a wake word associated with a specific voice assistant service;

monitoring, via a second wake word engine, the sound data stream from the one or more microphones for a wake word associated with the media playback system;

generating a first wake-word event corresponding to a first voice input when the first wake-word engine detects sound data matching the wake word associated with the specific voice assistant service in a first portion of the sound data stream;

based on generating the first wake-word event, sending sound data representing the first voice input to the specific voice assistant for processing of the first voice input;

generating a second wake-word event corresponding to a second voice input when the second wake-word engine detects sound data matching the wake word associated with the media playback system in a second portion of the sound data stream;

determining that the second voice input includes sound data matching one or more playback commands; and

based on (i) generating the second wake-word event and (ii) determining that the second voice input includes sound data matching one or more playback commands, sending sound data representing the second voice input to a voice assistant associated with the media playback system for processing of the second voice input.

9. The method of claim 8 , wherein sending the sound data representing the second voice input to the voice assistant associated with the media playback system comprises streaming, via a network interface, the sound data representing the second voice input to one or more remote servers associated with the media playback system.

10. The method of claim 8 , further comprising:

receiving, from the voice assistant associated with the media playback system, at least one playback instruction corresponding to the one or more playback commands; and

performing the at least one playback instruction.

11. The method of claim 10 , wherein the voice assistant associated with the media playback system identifies specific media content based on the second voice input, and wherein the at least one playback instruction instructs the playback device to play back the specific media content.

12. The method of claim 8 , further comprising:

generating a third wake-word event corresponding to a third voice input when the second wake-word engine detects sound data matching the wake word associated with the media playback system in a third portion of the sound data stream;

determining that the third voice input includes sound data matching a particular playback command; and

based on (i) generating the third wake-word event and (ii) determining that the third voice input includes sound data matching the particular playback command, performing the particular playback command.

13. The method of claim 8 , further comprising:

generating a third wake-word event corresponding to a third voice input when the second wake-word engine detects sound data matching the wake word associated with the media playback system in a third portion of the sound data stream;

determining that the third voice input includes sound data matching one or more smart home commands; and

based on (i) generating the third wake-word event and (ii) determining that the third voice input includes sound data matching one or more smart home commands, causing one or more smart home devices to carry out at least one smart home instruction corresponding to the one or more smart home commands.

14. The method of claim 13 , wherein causing the one or more smart home devices to carry out at least one smart home instruction corresponding to the one or more smart home commands comprises:

sending sound data representing the third voice input to a voice assistant for processing of the third voice input.

15. A non-transitory computer-readable medium having instructions stored thereon that are executable by one or more processors to cause a playback device to perform functions comprising:

capturing audio from one or more microphones into one or more buffers as a sound data stream;

monitoring, via a first wake word engine, the sound data stream from the one or more microphones for a wake word associated with a specific voice assistant service;

monitoring, via a second wake word engine, the sound data stream from the one or more microphones for a wake word associated with a media playback system, wherein the media playback system comprises the playback device;

generating a first wake-word event corresponding to a first voice input when the first wake-word engine detects sound data matching the wake word associated with the specific voice assistant service in a first portion of the sound data stream;

based on generating the first wake-word event, sending sound data representing the first voice input to the specific voice assistant for processing of the first voice input;

generating a second wake-word event corresponding to a second voice input when the second wake-word engine detects sound data matching the wake word associated with the media playback system in a second portion of the sound data stream;

determining that the second voice input includes sound data matching one or more playback commands; and

based on (i) generating the second wake-word event and (ii) determining that the second voice input includes sound data matching one or more playback commands, sending sound data representing the second voice input to a voice assistant associated with the media playback system for processing of the second voice input.

16. The non-transitory computer-readable medium of claim 15 , wherein sending the sound data representing the second voice input to the voice assistant associated with the media playback system comprises streaming, via a network interface, the sound data representing the second voice input to one or more remote servers associated with the media playback system.

17. The non-transitory computer-readable medium of claim 15 , wherein the functions further comprise:

receiving, from the voice assistant associated with the media playback system, at least one playback instruction corresponding to the one or more playback commands; and

performing the at least one playback instruction.

18. The non-transitory computer-readable medium of claim 17 , wherein the voice assistant associated with the media playback system identifies specific media content based on the second voice input, and wherein the at least one playback instruction instructs the playback device to play back the specific media content.

19. The non-transitory computer-readable medium of claim 15 , wherein the functions further comprise:

generating a third wake-word event corresponding to a third voice input when the second wake-word engine detects sound data matching the wake word associated with the media playback system in a third portion of the sound data stream;

determining that the third voice input includes sound data matching a particular playback command; and

based on (i) generating the third wake-word event and (ii) determining that the third voice input includes sound data matching the particular playback command, performing the particular playback command.

20. The non-transitory computer-readable medium of claim 15 , wherein the functions further comprise:

generating a third wake-word event corresponding to a third voice input when the second wake-word engine detects sound data matching the wake word associated with the media playback system in a third portion of the sound data stream;

determining that the third voice input includes sound data matching one or more smart home commands; and

based on (i) generating the third wake-word event and (ii) determining that the third voice input includes sound data matching one or more smart home commands, causing one or more smart home devices to carry out at least one smart home instruction corresponding to the one or more smart home commands.

21. The non-transitory computer-readable medium of claim 20 , wherein causing the one or more smart home devices to carry out at least one smart home instruction corresponding to the one or more smart home commands comprises:

sending sound data representing the third voice input to a voice assistant for processing of the third voice input.

Assignments (2)
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 13, 2020
From: WILBERDING, DAYN
To: SONOS, INC.
Reel/Frame 051817/0362 →
Continuity (4)
Continuation 16437437 · Jun 11, 2019
Continuation 16173797 · Oct 29, 2018
Continuation 15229868 · Aug 5, 2016
Related Publication 20200184980A1 · Jun 11, 2020
Cited By (32)
US 12,192,713 US 12,210,801 US 12,211,490 US 12,217,748 US 12,217,765 US 12,230,291 US 12,231,859 US 12,236,932 US 12,277,368 US 12,279,096 US 12,288,558 US 12,314,633 US 12,322,390 US 12,340,802 US 12,360,734 US 12,374,334 US 12,375,052 US 12,387,716 US 12,424,220 US 12,438,977 US 12,462,802 US 12,505,832 US 12,513,466 US 12,513,479 US 12,518,755 US 12,518,756 US 12,578,779 US 12,579,978 US 12,626,717 US 12,640,148 US 12,699,543 US 12,711,962