Systems and methods for power-efficient keyword detection
Systems and methods for audio processing include capturing first sound data via at least one microphone of a network microphone device (NMD) and determining, via a voice activity detection process, that the first sound data does not include voice activity. The first sound data is stored in a buffer, and the NMD forgoes spatial processing of the first sound data. The NMD can capture second sound data and determine, via the voice activity process, that the second sound data includes voice activity. The NMD spatially processes the second sound data to produce filtered sound data. The NMD detects a wake word based on data in the buffer. After detecting the wake word, the NMD may determine an action to be performed based on the data in the buffer.
1 . A network microphone device (NMD) comprising:
a plurality of microphones;
a network interface;
one or more processors; and
a tangible, non-transitory, computer-readable medium storing instructions that, when executed by the one or more processors, cause the NMD to perform operations comprising:
while operating in a first, low-power stage:
capturing first sound data via only a first subset of the microphones, the first subset comprising fewer than all of the plurality of microphones;
determining, via a voice activity detection process, that the first sound data includes voice activity; and
in response to determining that the first sound data includes voice activity, transitioning from the first, low-power stage to a second, high-power stage;
while operating in the second, high-power stage:
capturing second sound data via each of the plurality of microphones, the second sound data including a voice input;
performing further voice processing of the second sound data to produce filtered sound data;
detecting, via a keyword engine, a keyword based on the filtered sound data; and
after detecting the keyword, performing an action based at least in part on the voice input.
2 . The NMD of claim 1 , wherein performing further processing of the second sound data to produce filtered sound data comprises spatially processing the second sound data via a spatial processor.
3 . The NMD of claim 1 , wherein performing further processing of the second sound data to produce filtered sound data comprises performing acoustic echo cancellation on the second sound data.
4 . The NMD of claim 1 , wherein operations further comprise:
transmitting, via the network interface, at least a portion of the second sound data to one or more remote computing devices associated with a voice assistant service; and
receiving, via the network interface, a determined intent based on the at least a portion of the second sound data.
5 . The NMD of claim 1 , wherein the action comprises controlling playback of audio content.
6 . The NMD of claim 1 , wherein, in the first stage, the one or more processors operate at a lower average clock speed than in the second stage.
7 . The NMD of claim 1 , wherein the operations further comprise, after performing the action, transitioning the NMD from the second stage to the first stage.
8 . A method comprising:
while operating in a first, low-power stage:
capturing first sound data via only a first subset of microphones of a network microphone device (NMD), the first subset comprising fewer than all of the microphones of the NMD;
determining, via a voice activity detection process of the NMD, that the first sound data includes voice activity; and
in response to determining that the first sound data includes voice activity, transitioning from the first, low-power stage to a second, high-power stage;
while operating in the second, high-power stage:
capturing second sound data via each of the microphones of the NMD, the second sound data including a voice input;
performing further voice processing of the second sound data to produce filtered sound data;
detecting, via a keyword engine of the NMD, a keyword based on filtered sound data; and
after detecting the keyword, performing an action based at least in part on the voice input.
9 . The method of claim 8 , wherein performing further processing of the second sound data to produce filtered sound data comprises spatially processing the second sound data via a spatial processor.
10 . The method of claim 8 , wherein performing further processing of the second sound data to produce filtered sound data comprises performing acoustic echo cancellation on the second sound data.
11 . The method of claim 9 , further comprising:
transmitting, via a network interface of the NMD, at least a portion of the second sound data to one or more remote computing devices associated with a voice assistant service; and
receiving, via the network interface, a determined intent based on the at least a portion of the second sound data.
12 . The method of claim 8 , wherein the action comprises controlling playback of audio content.
13 . The method of claim 8 , wherein, in the first stage, one or more processors of the NMD operates at a lower average clock speed than in the second stage.
14 . The method of claim 8 , further comprising, after performing the action, transitioning the NMD from the second stage to the first stage.
15 . One or more tangible, non-transitory computer-readable media storing instructions that, when executed by one or more processors of a network microphone device (NMD), cause the NMD to perform operations comprising:
while operating in a first, low-power stage:
capturing first sound data via only a first subset of microphones of the NMD, the first subset comprising fewer than all of the microphones of the NMD;
determining, via a voice activity detection process of the NMD, that the first sound data includes voice activity; and
in response to determining that the first sound data includes voice activity, transitioning from the first, low-power stage to a second, high-power stage;
while operating in the second, high-power stage:
capturing second sound data via each of the microphones of the NMD, the second sound data including a voice input;
performing further voice processing of the second sound data to produce filtered sound data;
detecting, via a keyword engine, a keyword based on the filtered sound data; and
after detecting the keyword, performing an action based at least in part on the voice input.
16 . The one or more computer-readable media of claim 15 , wherein performing further processing of the second sound data to produce filtered sound data comprises spatially processing the second sound data via a spatial processor.
17 . The one or more computer-readable media of claim 15 , wherein performing further processing of the second sound data to produce filtered sound data comprises performing acoustic echo cancellation on the second sound data.
18 . The one or more computer-readable media of claim 15 , wherein the operations further comprise:
transmitting, via a network interface of the NMD, at least a portion of the second sound data to one or more remote computing devices associated with a voice assistant service; and
receiving, via the network interface, a determined intent based on the at least a portion of the second sound data.
19 . The one or more computer-readable media of claim 15 , wherein the action comprises controlling playback of audio content.
20 . The one or more computer-readable media of claim 15 , wherein the operations further comprise:
after performing the action, transitioning the NMD from the second stage to the first stage.