IP Library Granted Patent US 11,380,322
Granted Patent B2
US 11,380,322 · App. 16/679,538 · Granted Jul 5, 2022

Wake-word detection suppression

Inventor: Jonathan P. Lang (Santa Barbara, CA)
Assignee: Sonos, Inc.
G10L15/22G06F3/165G06F3/167H04N21/42203G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,380,322
App. No.
16/679,538
Granted
Jul 5, 2022
Kind
B2
Abstract

Example techniques involve suppressing a wake word response to a local wake word. An example implementation involves a playback device receiving audio content for playback by the playback device and providing a sound data stream representing the received audio content to a voice assistant service (VAS) wake-word engine and a local keyword engine. The playback device plays back a first portion of the audio content and detects, via the local keyword engine, that a second portion of the received audio content includes sound data matching one or more particular local keywords. Before the second portion of the received audio content is played back, the playback device disables a local keyword response of the local keyword engine to the one or more particular local keywords and then plays back the second portion of the audio content via one or more speakers.

Claims (70)

1. A playback device comprising:

a network interface;

one or more microphones;

one or more processors;

data storage having stored therein instructions executable by the one or more processors to cause the playback device to perform functions comprising:

receiving audio content for playback by the playback device;

providing a sound data stream representing the received audio content to (i) a voice assistant service (VAS) wake-word engine and (ii) a local wakeword engine, wherein the VAS wake-word engine is operable to (a) generate a VAS wake word response when the VAS wake-word engine detects a VAS wake word in a microphone sound data stream representing sound detected by one or more microphones of the playback device and (b) stream sound data representing the sound detected by the one or more microphones to one or more servers of the VAS when the VAS wake word response is generated, and wherein the local wakeword engine is operable to (a) generate a local wakeword response when the local wakeword engine detects one or more local wakewords in the microphone sound data stream representing sound detected by the one or more microphones and (b) determine an intent of a voice input comprising the one or more local wakewords when the local wakeword response is generated;

playing back a first portion of the audio content via one or more speakers;

detecting, via the local wakeword engine, that a second portion of the received audio content includes sound data matching one or more particular local wakewords;

before the second portion of the received audio content that includes the sound data matching the one or more particular local wakewords is played back, disabling the local wakeword response of the local wakeword engine to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device; and

playing back the second portion of the audio content via one or more speakers.

2. The playback device of claim 1 , wherein the playback device is connected via a local area network to one or more networked microphone devices comprising respective local wakeword engines, and wherein the functions further comprise:

causing the one or more networked microphone devices to disable their respective local wakeword responses to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device.

3. The playback device of claim 2 , wherein causing the one or more networked microphone devices to disable their respective local wakeword responses to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device comprises:

sending, via the network interface to the one or more networked microphone devices, instructions that cause the one or more networked microphone devices to disable their respective local wakeword responses to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device.

4. The playback device of claim 3 , wherein the one or more networked microphones devices are a subset of networked microphone devices connected to the local area network, and wherein the functions further comprise:

determining that the one or more networked microphone devices are in audible vicinity of the audio content; and

sending the instructions that cause the one or more networked microphone devices to disable their respective local wakeword responses to the one or more particular local wakewords during playback of the audio content by the playback device based on determining that the one or more networked microphone devices are in audible vicinity of the audio content.

5. The playback device of claim 1 , wherein disabling the local wakeword response of the local wakeword engine to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device comprises:

before playing back the second portion of the audio content, modifying the second portion of the audio content to incorporate acoustic markers in segments of the second portion that represent respective local wakewords, wherein the acoustic markers causes the local wakeword engine to disable its local wakeword responses to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device.

6. The playback device of claim 1 , wherein the functions further comprise:

detecting, via the VAS wake-word engine, that a third portion of the received audio content includes sound data matching a particular VAS wake word;

before the third portion of the received audio content that includes the sound data matching the particular VAS wake word is played back, disabling the VAS wake word response of the VAS wake-word engine to the particular VAS wake word during playback of the third portion of the audio content by the playback device; and

playing back the third portion of the audio content via one or more speakers.

7. The playback device of claim 1 , wherein the VAS wake-word engine comprises a first VAS wake-word detection algorithm for a first VAS and a second wake-word detection algorithm for a second VAS, and wherein providing the sound data stream representing the received audio content to the VAS wake-word engine comprises:

applying, to the sound data stream representing the received audio content before the audio content is played back by the playback device, the first VAS wake word detection algorithm for the first VAS; and

applying, to the sound data stream representing the received audio content before the audio content is played back by the playback device, the second VAS wake word detection algorithm for the second VAS.

8. A method to be performed by a playback device, the method comprising:

receiving audio content for playback by the playback device;

providing a sound data stream representing the received audio content to (i) a voice assistant service (VAS) wake-word engine and (ii) a local wakeword engine, wherein the VAS wake-word engine is operable to (a) generate a VAS wake word response when the VAS wake-word engine detects a VAS wake word in a microphone sound data stream representing sound detected by one or more microphones of the playback device and (b) stream sound data representing the sound detected by the one or more microphones to one or more servers of the VAS when the VAS wake word response is generated, and wherein the local wakeword engine is operable to (a) generate a local wakeword response when the local wakeword engine detects one or more local wakewords in the microphone sound data stream representing sound detected by the one or more microphones and (b) determine an intent of a voice input comprising the one or more local wakewords when the local wakeword response is generated;

playing back a first portion of the audio content via one or more speakers;

detecting, via the local wakeword engine, that a second portion of the received audio content includes sound data matching one or more particular local keywords;

before the second portion of the received audio content that includes the sound data matching the one or more particular local wakewords is played back, disabling the local wakeword response of the local wakeword engine to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device; and

playing back the second portion of the audio content via one or more speakers.

9. The method of claim 8 , wherein the playback device is connected via a local area network to one or more networked microphone devices comprising respective local wakeword engines, and wherein the method further comprises:

causing the one or more networked microphone devices to disable their respective local keyword responses to the one or more particular local keywords during playback of the second portion of the audio content by the playback device.

10. The method of claim 9 , wherein causing the one or more networked microphone devices to disable their respective local wakeword responses to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device comprises:

sending, via a network interface to the one or more networked microphone devices, instructions that cause the one or more networked microphone devices to disable their respective local wakeword responses to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device.

11. The method of claim 10 , wherein the one or more networked microphones devices are a subset of networked microphone devices connected to the local area network, and wherein the method further comprises:

determining that the one or more networked microphone devices are in audible vicinity of the audio content; and

sending the instructions that cause the one or more networked microphone devices to disable their respective local wakeword responses to the one or more particular local wakewords during playback of the audio content by the playback device based on determining that the one or more networked microphone devices are in audible vicinity of the audio content.

12. The method of claim 8 , wherein disabling the local wakeword response of the local wakeword engine to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device comprises:

before playing back the second portion of the audio content, modifying the second portion of the audio content to incorporate acoustic markers in segments of the second portion that represent respective local wakewords, wherein the acoustic markers causes the local wakeword engine to disable its local wakeword responses to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device.

13. The method of claim 8 , further comprising:

detecting, via the VAS wake-word engine, that a third portion of the received audio content includes sound data matching a particular VAS wake word;

before the third portion of the received audio content that includes the sound data matching the particular VAS wake word is played back, disabling the VAS wake word response of the VAS wake-word engine to the particular VAS wake word during playback of the third portion of the audio content by the playback device; and

playing back the third portion of the audio content via one or more speakers.

14. The method of claim 8 , wherein the VAS wake-word engine comprises a first VAS wake-word detection algorithm for a first VAS and a second wake-word detection algorithm for a second VAS, and wherein providing the sound data stream representing the received audio content to the VAS wake-word engine comprises:

applying, to the sound data stream representing the received audio content before the audio content is played back by the playback device, the first VAS wake word detection algorithm for the first VAS; and

applying, to the sound data stream representing the received audio content before the audio content is played back by the playback device, the second VAS wake word detection algorithm for the second VAS.

15. A tangible, non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors of a playback device, cause the playback device to perform functions comprising:

receiving audio content for playback by the playback device;

providing a sound data stream representing the received audio content to (i) a voice assistant service (VAS) wake-word engine and (ii) a local wakeword engine, wherein the VAS wake-word engine is operable to (a) generate a VAS wake word response when the VAS wake-word engine detects a VAS wake word in a microphone sound data stream representing sound detected by one or more microphones of the playback device and (b) stream sound data representing the sound detected by the one or more microphones to one or more servers of the VAS when the VAS wake word response is generated, and wherein the local wakeword engine is operable to (a) generate a local wakeword response when the local wakeword engine detects one or more local wakewords in the microphone sound data stream representing sound detected by the one or more microphones and (b) determine an intent of a voice input comprising the one or more local wakewords when the local wakeword response is generated;

playing back a first portion of the audio content via one or more speakers;

detecting, via the local wakeword engine, that a second portion of the received audio content includes sound data matching one or more particular local wakewords;

before the second portion of the received audio content that includes the sound data matching the one or more particular local wakewords is played back, disabling the local wakeword response of the local wakeword engine to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device; and

playing back the second portion of the audio content via one or more speakers.

16. The tangible, non-transitory computer-readable medium of claim 15 , wherein the playback device is connected via a local area network to one or more networked microphone devices comprising respective local wakeword engines, and wherein the functions further comprise:

causing the one or more networked microphone devices to disable their respective local wakeword responses to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device.

17. The tangible, non-transitory computer-readable medium of claim 16 , wherein causing the one or more networked microphone devices to disable their respective local wakeword responses to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device comprises:

sending, via a network interface to the one or more networked microphone devices, instructions that cause the one or more networked microphone devices to disable their respective local wakeword responses to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device.

18. The tangible, non-transitory computer-readable medium of claim 17 , wherein the one or more networked microphones devices are a subset of networked microphone devices connected to the local area network, and wherein the functions further comprise:

determining that the one or more networked microphone devices are in audible vicinity of the audio content; and

sending the instructions that cause the one or more networked microphone devices to disable their respective local wakeword responses to the one or more particular local wakewords during playback of the audio content by the playback device based on determining that the one or more networked microphone devices are in audible vicinity of the audio content.

19. The tangible, non-transitory computer-readable medium of claim 15 , wherein disabling the local wakeword response of the local wakeword engine to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device comprises:

before playing back the second portion of the audio content, modifying the second portion of the audio content to incorporate acoustic markers in segments of the second portion that represent respective local wakewords, wherein the acoustic markers causes the local wakeword engine to disable its local wakeword responses to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device.

20. The tangible, non-transitory computer-readable medium of claim 15 , wherein the functions further comprise:

detecting, via the VAS wake-word engine, that a third portion of the received audio content includes sound data matching a particular VAS wake word;

before the third portion of the received audio content that includes the sound data matching the particular VAS wake word is played back, disabling the VAS wake word response of the VAS wake-word engine to the particular VAS wake word during playback of the third portion of the audio content by the playback device; and

playing back the third portion of the audio content via one or more speakers.

Assignments (2)
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 11, 2019
From: LANG, JONATHAN P.
To: SONOS, INC.
Reel/Frame 050969/0651 →
Continuity (2)
Continuation 15670361 · Aug 7, 2017
Related Publication 20200075010A1 · Mar 5, 2020