IP Library › Granted Patent US 11,551,700
Granted Patent B2
US 11,551,700 · App. 17/248,427 · Granted Jan 10, 2023

Systems and methods for power-efficient keyword detection

Inventors: Aaron Jones (San Diego, CA); Saeed Bagheri Sereshki (Goleta, CA); Daniele Giacobello (Los Angeles, CA)
Assignee: Sonos, Inc.
G10L17/22G10L15/02G10L15/05G10L15/22G10L17/02G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,551,700
App. No.
17/248,427
Granted
Jan 10, 2023
Kind
B2
Abstract

Systems and methods for audio processing include capturing sound data via at least one microphone of a network microphone device (NMD) and determining whether the captured sound includes voice activity. While in a first stage, the NMD forgoes spatial processing of the captured sound data. If the NMD determines that the detected sound includes voice activity, the NMD transitions to a second stage. In this second stage, the NMD spatially processes the detected sound to produce filtered sound data and detects a wake word. After detecting the wake word, the NMD may determine an action to be performed based on the captured sound data.

Claims (56)

1. A network microphone device (NMD) comprising:

a plurality of microphones;

a network interface;

one or more processors; and

a tangible, non-transitory, computer-readable medium storing instructions that, when executed by the one or more processors, cause the NMD to perform operations comprising:

detecting sound via at least one of the microphones;

determining, via a voice activity detection process, whether the detected sound includes voice activity;

capturing the detected sound as sound data;

while in a first stage, forgoing spatial processing of the captured sound data;

transitioning the NMD from the first stage to a second stage after determining that the detected sound includes voice activity; and

while in the second stage:

spatially processing the captured sound data using a spatial processor to produce filtered sound data;

detecting, via a wake-word engine, a wake word based on captured the filtered sound data; and

after detecting the wake word, determining an action to be performed based on the captured sound data.

2. The NMD of claim 1 , wherein the captured sound data comprises less than about 300 milliseconds of sound data.

3. The NMD of claim 1 , wherein determining the action to be performed based on the captured sound data comprises:

transmitting, via the network interface, at least a portion of the captured sound data to one or more remote computing devices associated with a voice assistant service corresponding to the detected wake word for processing a voice utterance.

4. The NMD of claim 1 , wherein the operations further comprise performing the determined action, wherein the action comprises controlling playback of audio content.

5. The NMD of claim 1 , wherein, in the first stage, the one or more processors operate at a lower average clock speed than in the second stage.

6. The NMD of claim 1 , wherein, in the first stage, the NMD consumes less power than in the second stage.

7. The NMD of claim 1 , wherein, in the first stage, the spatial processor is in a low-power standby mode.

8. The NMD of claim 1 , wherein the operations further comprise:

after determining the action to be performed, transitioning the NMD from the second stage to the first stage.

9. A method comprising:

detecting sound via at least one microphone of a network microphone device (NMD);

determining, via a voice activity detection process of the NMD, whether the detected sound includes voice activity;

capturing the detected sound as sound data;

while in a first stage of the NMD, forgoing spatial processing of the captured sound data;

transitioning the NMD from the first stage to a second stage after determining that the detected sound includes voice activity; and

while in the second stage:

spatially processing the captured sound data using a spatial processor to produce filtered sound data;

detecting, via a wake-word engine, a wake word based on the filtered sound data; and

after detecting the wake word, determining an action to be performed based on the captured sound data.

10. The method of claim 9 , wherein the captured sound data comprises less than about 300 milliseconds of sound data.

11. The method of claim 9 , wherein determining the action to be performed based on the captured sound data comprises:

transmitting, via a network interface of the NMD, at least a portion of the captured sound data to one or more remote computing devices associated with a voice assistant service corresponding to the detected wake word for processing a voice utterance.

12. The method of claim 9 , further comprising performing the determined action, wherein the action comprises controlling playback of audio content.

13. The method of claim 9 , wherein, in the first stage, the NMD consumes less power than in the second stage.

14. The method of claim 9 , further comprising, after determining the action to be performed, transitioning the NMD from the second stage to the first stage.

15. A tangible, non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a network microphone device (NMD), cause the NMD to perform operations comprising:

detecting sound via at least one microphone of the NMD;

determining, via a voice activity detection process, whether the detected sound includes voice activity;

capturing the detected sound as sound data;

while in a first stage, forgoing spatial processing of the captured sound data;

transitioning the NMD from the first stage to a second stage after determining that the detected sound includes voice activity; and

while in the second stage:

spatially processing the captured sound data using a spatial processor to produce filtered sound data;

detecting, via a wake-word engine, a wake word based on the filtered sound data; and

after detecting the wake word, determining an action to be performed based on the captured sound data.

16. The computer-readable medium of claim 15 , wherein the captured sound data comprises less than about 300 milliseconds of sound data.

17. The computer-readable medium of claim 15 , wherein determining the action to be performed based on the captured sound data comprises:

transmitting, via a network interface of the NMD, at least a portion of the captured sound data to one or more remote computing devices associated with a voice assistant service corresponding to the detected wake word for processing a voice utterance.

18. The computer-readable medium of claim 15 , wherein the operations further comprise performing the determined action, wherein the action comprises controlling playback of audio content.

19. The computer-readable medium of claim 15 , wherein, in the first stage, the NMD consumes less power than in the second stage.

20. The computer-readable medium of claim 15 , wherein the operations further comprise:

after determining the action to be performed, transitioning the NMD from the second stage to the first stage.

Assignments (2)
SECURITY INTEREST Recorded Jan 30, 2026
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074533/0615 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2021
From: JONES, AARON; SERESHKI, SAEED BAGHERI; GIACOBELLO, DANIELE
To: SONOS, INC.
Reel/Frame 055194/0563 →
Continuity (1)
Related Publication 20220238120A1 · Jul 28, 2022
Cited By (1)
US 12,749,486