IP Library Granted Patent US 11,557,294
Granted Patent B2
US 11,557,294 · App. 17/454,676 · Granted Jan 17, 2023

Systems and methods of operating media playback systems having multiple voice assistant services

Inventors: Ryan Richard Myers (Santa Barbara, CA); Luis R. Vega Zayas (Arlington, MA); Sangah Park (Somerville, MA)
Assignee: Sonos, Inc.
G10L15/22G06F3/165G10L15/08G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,557,294
App. No.
17/454,676
Granted
Jan 17, 2023
Kind
B2
Abstract

Systems and methods for managing multiple voice assistants are disclosed. Audio input is received via one or more microphones of a playback device. A first activation word is detected in the audio input via the playback device. After detecting the first activation word, the playback device transmits a voice utterance of the audio input to a first voice assistant service (VAS). The playback device receives, from the first VAS, first content to be played back via the playback device. The playback device also receives, from a second VAS, second content to be played back via the playback device. The playback device plays back the first content while suppressing the second content. Such suppression can include delaying or canceling playback of the second content.

Claims (73)

1. A method comprising:

receiving an audio input via one or more microphones of a playback device;

monitoring the audio input for a first activation word associated with a first voice assistant service (VAS) and for a second activation word associated with a second VAS different from the first VAS;

detecting, via a first activation-word detector of the playback device, the first activation word in the audio input;

after detecting the first activation word, suppressing monitoring the audio input for the second activation word;

after detecting the first activation word, transmitting, via the playback device, a voice utterance of the audio input to one or more remote computing devices associated with the first VAS;

receiving, from the one or more remote computing devices associated with the first VAS, first content to be played back via the playback device;

playing back, via the playback device, the first content;

while playing back the first content, receiving, from one or more remote computing device associated with the second VAS, second content to be played back via the playback device;

temporarily reducing a volume of playback of the first content;

while the volume is temporarily reduced, playing back, via the playback device, the second content;

after playing back the second content, restoring the volume of playback of the first content; and

after restoring the volume of playback of the first content, resuming monitoring the audio input for at least the second activation word.

2. The method of claim 1 , further comprising:

after receiving the second content, arbitrating between the first content and the second content; and

based on the arbitration, temporarily reducing the volume of playback of the first content while playing back the second content.

3. The method of claim 2 , wherein the arbitrating is based at least on a characteristic of at least one of the first content or the second content.

4. The method of claim 3 , wherein the characteristics of the first and second contents considered in the arbitrating step comprises:

the first content comprises a text-to-speech output; and

the second content comprises at least one of: an alarm, a user broadcast, or a text-to-speech output.

5. The method of claim 3 , wherein the first and second content have the same category of content.

6. The method of claim 3 , wherein the second content is one of: a timer or an alarm.

7. The method of claim 1 , further comprising suppressing monitoring audio input for the second activation word associated with the second VAS while a user is interacting with the first VAS.

8. A playback device comprising:

one or more microphones;

a network interface;

one or more audio transducers;

one or more processors; and

data storage having instructions stored thereon that, when executed by the one or more processors, cause the playback device to perform operations comprising:

receiving an audio input via the one or more microphones;

monitoring the audio input for a first activation word associated with a first voice assistant service (VAS) and for a second activation word associated with a second VAS different from the first VAS;

detecting, via a first activation-word detector of the playback device, the first activation word in the audio input;

after detecting the first activation word, suppressing monitoring the audio input for the second activation word;

after detecting the first activation word, transmitting, via the network interface, a voice utterance of the audio input to one or more remote computing devices associated with the first VAS;

receiving, from the one or more remote computing devices associated with the first VAS, first content to be played back via the playback device;

playing back, via the one or more audio transducers of the playback device, the first content;

while playing back the first content, receiving, from one or more remote computing device associated with the second VAS, second content to be played back via the playback device;

temporarily reducing a volume of playback of the first content;

while the volume is temporarily reduced, playing back, via the playback device, the second content;

after playing back the second content, restoring the volume of playback of the first content; and

after restoring the volume of playback of the first content, resuming monitoring the audio input for at least the second activation word.

9. The playback device of claim 8 , wherein the operations further comprise:

after receiving the second content, arbitrating between the first content and the second content; and

based on the arbitration, temporarily reducing the volume of playback of the first content while playing back the second content.

10. The playback device of claim 9 , wherein the arbitrating is based at least on a characteristic of at least one of the first content or the second content.

11. The playback device of claim 10 , wherein the characteristics of the first and second contents considered in the arbitrating step comprises:

the first content comprises a text-to-speech output; and

the second content comprises at least one of: an alarm, a user broadcast, or a text-to-speech output.

12. The playback device of claim 10 , wherein the first and second content have the same category of content.

13. The playback device of claim 10 , wherein the second content is one of: a timer or an alarm.

14. The playback device of claim 8 , wherein the operations further comprise suppressing monitoring audio input for the second activation word associated with the second VAS while a user is interacting with the first VAS.

15. A tangible, non-transitory computer readable media storing instructions that, when executed by one or more processors of a playback device, cause the playback device to perform operations comprising:

receiving an audio input via one or more microphones of the playback device;

monitoring the audio input for a first activation word associated with a first voice assistant service (VAS) and for a second activation word associated with a second VAS different from the first VAS;

detecting, via a first activation-word detector of the playback device, the first activation word in the audio input;

after detecting the first activation word, suppressing monitoring the audio input for the second activation word;

after detecting the first activation word, transmitting, via a network interface of the playback device, a voice utterance of the audio input to one or more remote computing devices associated with the first VAS;

receiving, from the one or more remote computing devices associated with the first VAS, first content to be played back via the playback device;

playing back, via one or more audio transducers of the playback device, the first content;

while playing back the first content, receiving, from one or more remote computing device associated with the second VAS, second content to be played back via the playback device;

temporarily reducing a volume of playback of the first content;

while the volume is temporarily reduced, playing back, via the playback device, the second content;

after playing back the second content, restoring the volume of playback of the first content; and

after restoring the volume of playback of the first content, resuming monitoring the audio input for at least the second activation word.

16. The computer-readable media of claim 15 , wherein the operations further comprise:

after receiving the second content, arbitrating between the first content and the second content; and

based on the arbitration, temporarily reducing the volume of playback of the first content while playing back the second content.

17. The computer-readable media of claim 16 , wherein the arbitrating is based at least on a characteristic of at least one of the first content or the second content.

18. The computer-readable media of claim 17 , wherein the characteristics of the first and second contents considered in the arbitrating step comprises:

the first content comprises a text-to-speech output; and

the second content comprises at least one of: an alarm, a user broadcast, or a text-to-speech output.

19. The computer-readable media claim 17 , wherein the first and second content have the same category of content.

20. The computer-readable media of claim 17 , wherein the second content is one of: a timer or an alarm.

Assignments (2)
SECURITY INTEREST Recorded Jan 30, 2026
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074533/0615 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2021
From: VEGA ZAYAS, LUIS R.; MYERS, RYAN RICHARD; PARK, SANGAH
To: SONOS, INC.
Reel/Frame 058098/0815 →
Continuity (2)
Continuation 16213570 · Dec 7, 2018
Related Publication 20220076675A1 · Mar 10, 2022
Cited By (1)
US 12,284,417