IP Library › Granted Patent US 11,100,923
Granted Patent B2
US 11,100,923 · App. 16/145,275 · Granted Aug 24, 2021

Systems and methods for selective wake word detection using neural network models

Inventors: Joachim Fainberg (Oslo, NO); Daniele Giacobello (Los Angeles, CA); Klaus Hartung (Santa Barbara, CA)
Assignee: Sonos, Inc.
G10L15/22G10L15/14G10L15/16G10L15/30G10L15/32G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,100,923
App. No.
16/145,275
Granted
Aug 24, 2021
Kind
B2
Abstract

Systems and methods for media playback via a media playback system include capturing sound data via a network microphone device and identifying a candidate wake word in the sound data. Based on identification of the candidate wake word in the sound data, the system selects a first wake-word engine from a plurality of wake-word engines. Via the first wake-word engine, the system analyzes the sound data to detect a confirmed wake word, and, in response to detecting the confirmed wake word, transmits a voice utterance of the sound data to one or more remote computing devices associated with a voice assistant service.

Claims (44)

1. A method comprising:

capturing sound data via a network microphone device;

identifying, via the network microphone device, using a keyword spotting algorithm, a candidate wake word in the sound data, the network microphone device comprising a plurality of wake-world engines in a low-power or no-power state;

based on identification of the candidate wake word in the sound data, selecting a first wake-word engine from the plurality of wake-word engines, wherein the first wake-word engine is associated with a first voice assistant service and another of the plurality of wake-word engines is associated with a second voice assistant service different from the first, wherein selecting the first wake-word engine comprises activating the first wake-word engine from the low-power or no-power state to a high-power state, while the other(s) of the plurality of wake-word engines remain(s) in the low-power or no-power state;

with the first wake-word engine, analyzing the sound data to confirm detection of a wake word, wherein the first wake-word engine is configured to determine whether the candidate wake work is present in the sound data with a higher accuracy than the keyword spotting algorithm; and

in response to confirming the detection of the wake word, transmitting a voice utterance of the sound data to one or more remote computing devices associated with the first voice assistant service.

2. The method of claim 1 , wherein identifying the candidate wake word comprises determining a probability that the candidate wake word is present in the sound data.

3. The method of claim 1 , wherein the first wake-word engine is associated with the candidate wake word, and wherein another of the plurality of wake-word engines is associated with one or more additional wake words.

4. The method of claim 1 , wherein identifying the candidate wake word comprises applying a neural network model to the sound data.

5. The method of claim 4 , wherein the neural network model comprises a compressed neural network model stored locally on the network microphone device.

6. The method of claim 1 , further comprising, after transmitting the voice utterance, receiving, via the network microphone device, a selection of media content related to the voice utterance.

7. The method of claim 1 , wherein the plurality of wake-word engines comprises:

the first wake-word engine; and

a second wake-word engine configured to perform a local function of the network microphone device.

8. A network microphone device, comprising:

one or more processors;

at least one microphone; and

tangible, non-transitory, computer-readable media storing instructions executable by one or more processors to cause the network microphone device to perform operations comprising:

capturing sound data via the network microphone device;

identifying, via the network microphone device, using a keyword spotting algorithm, a candidate wake word in the sound data, the network microphone device comprising a plurality of wake-word engines in a low-power or no-power state;

based on identification of the candidate wake word in the sound data, selecting a first wake-word engine from the plurality of wake-word engines, wherein the first wake-word engine is associated with a first voice assistant service and another of the plurality of wake-word engines is associated with a second voice assistant service different from the first, and wherein selecting the first wake-word engine comprises activating the first wake-word engine from the low-power or no-power state to a high-power state while the other(s) of the plurality of wake-word engines remain(s) in the low-power or no-power state;

with the first wake-word engine, analyzing the sound data to confirm detection of a wake word, wherein the first wake-word engine is configured to determine whether the candidate wake word is present in the sound data with a higher accuracy than the keyword spotting algorithm; and

in response to confirming detection of the wake word, transmitting a voice utterance of the sound data to one or more remote computing devices associated with the first voice assistant service.

9. The network microphone device of claim 8 , wherein identifying the candidate wake word comprises determining a probability that the candidate wake word is present in the sound data.

10. The network microphone device of claim 8 , wherein the first wake-word engine is associated with the candidate wake word, and wherein another of the plurality of wake-word engines is associated with one or more additional wake words.

11. The network microphone device of claim 8 , wherein identifying the candidate wake word comprises applying a neural network model to the sound data.

12. The network microphone device of claim 11 , wherein the neural network model comprises a compressed neural network model stored locally on the network microphone device.

13. The network microphone device of claim 8 , wherein the operations further comprise, after transmitting the voice utterance, receiving, via the network microphone device, a selection of media content related to the voice utterance.

14. The network microphone device of claim 8 , wherein the plurality of wake-word engines comprises:

the first wake-word engine; and

a second wake-word engine configured to perform a local function of the network microphone device.

15. Tangible, non-transitory, computer-readable media storing instructions executable by one or more processors to cause a network microphone device to perform operations comprising:

capturing sound data via the network microphone device;

identifying, via the network microphone device, using a keyword spotting algorithm, a candidate wake word in the sound data, the network microphone device comprising a plurality of wake-word engines in a low-power or no-power state;

based on identification of the candidate wake word in the sound data, selecting a first wake-word engine from the plurality of wake-word engines, wherein the first wake-word engine is associated with a first voice assistant service and another of the plurality of wake-word engines is associated with a second voice assistant service different from the first, and wherein selecting the first wake-word engine comprises activating the first wake-word engine from the low-power or no-power state to a high-power state while the other(s) of the plurality of wake-word engines remain(s) in the low-power or no-power state;

with the first wake-word engine, analyzing the sound data to confirm detection of a wake word, wherein the first wake-word engine is configured to determine whether the candidate wake work is present in the sound data with a higher accuracy than the keyword spotting algorithm; and

in response to confirming detection of the wake word, transmitting a voice utterance of the sound data to one or more remote computing devices associated with the first voice assistant service.

16. The tangible, non-transitory, computer-readable media of claim 15 , wherein identifying the candidate wake word comprises determining a probability that the candidate wake word is present in the sound data.

17. The tangible, non-transitory, computer-readable media of claim 15 , wherein the first wake-word engine is associated with the candidate wake word, and wherein another of the plurality of wake-word engines is associated with one or more additional wake words.

18. The tangible, non-transitory, computer-readable media of claim 15 , wherein identifying the candidate wake word comprises applying a locally stored neural network model to the sound data.

19. The tangible, non-transitory, computer-readable media of claim 15 , wherein the operations further comprise, after transmitting the voice utterance, receiving, via the network microphone device, a selection of media content related to the voice utterance.

20. The tangible, non-transitory, computer-readable media of claim 15 , wherein the plurality of wake-word engines comprises:

the first wake-word engine; and

a second wake-word engine configured to perform a local function of the network microphone device.

Assignments (2)
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2018
From: FAINBERG, JOACHIM; GIACOBELLO, DANIELE; HARTUNG, KLAUS
To: SONOS, INC.
Reel/Frame 047483/0248 →
Continuity (1)
Related Publication 20200105256A1 · Apr 2, 2020
Cited By (34)
US 12,192,713 US 12,211,490 US 12,212,945 US 12,217,748 US 12,230,266 US 12,230,291 US 12,236,932 US 12,277,368 US 12,279,096 US 12,283,269 US 12,288,558 US 12,322,390 US 12,327,549 US 12,327,556 US 12,360,734 US 12,375,052 US 12,387,716 US 12,424,220 US 12,438,977 US 12,505,832 US 12,513,466 US 12,513,479 US 12,518,755 US 12,518,756 US 12,579,978 US 12,597,425 US 12,626,717 US 12,640,148 US 12,694,278 US 12,699,543 US 12,711,962 US 12,732,547 US 12,744,035 US 12,748,566