Methods and systems for voice control
One or more portions of audio input may be detected. Timing data associated with the one or more portions of audio may be determined. Audio processing may be carried out based on the timing data.
1 . A method comprising:
receiving, a first portion of audio data and a second portion of audio data;
determining, based on the first portion of audio data and the second portion of audio data, a first wake word confidence score;
determining, based on the second portion of audio data, a second wake word confidence score;
determining, based on the second wake word confidence score being less than the first wake word confidence score, that the first portion of audio data comprises a beginning of a wake word; and
based on the first portion of audio data comprising the beginning of the wake word, sending the first portion of audio data for processing and excluding the second portion of audio data.
2 . The method of claim 1 , wherein the first portion of audio data and the second portion of audio data comprise one or more user utterances received by a voice enabled device.
3 . The method of claim 1 , wherein the first wake word confidence score and the second wake word confidence score are configured to indicate the first portion of audio data or the second portion of audio data comprises a portion of the wake word.
4 . The method of claim 1 , further comprising opening a communication channel with a voice service based on the second wake word confidence score being less than the first wake word confidence score.
5 . The method of claim 1 , wherein sending the first portion of audio data and the second portion of audio data for processing comprises sending the first portion of audio data and the second portion of audio data to a cloud based voice service.
6 . The method of claim 1 , further comprising determining, based on the second wake word confidence score being less than the first wake word confidence score, that the second portion of audio data comprises an end of the wake word.
7 . The method of claim 1 , further comprising determining a wake word duration.
8 . A method comprising:
receiving a plurality of portions of audio data comprising a wake word;
determining, based on a first portion of audio data of the plurality of portions of audio data and a second portion of audio data of the plurality of portions of audio data, a first wake word confidence score;
determining, based on the second portion of audio data and a third portion of audio data of the plurality of portions of audio data, a second wake word confidence score;
determining, based on the second wake word confidence score being higher than the first wake word confidence score, that a time stamp associated with the second portion of audio data corresponds to a beginning of the wake word; and
based on determining the time stamp associated with the second portion of audio data corresponds to the beginning of the wake word, opening a communication channel with a service.
9 . The method of claim 8 , wherein the plurality of portions of audio data comprise one or more voice inputs received from a voice enabled device.
10 . The method of claim 8 , wherein the first wake word confidence score indicates a likelihood the first portion of audio data or the second portion of audio data comprise a wake word.
11 . The method of claim 8 , wherein determining that the time stamp associated with the second portion of audio data corresponds to the beginning of the wake word comprises determining an audio analysis window.
12 . The method of claim 8 , further comprising processing the plurality of portions of audio data, wherein processing the plurality of portions of audio data comprises one or more of: responding to a query, sending a message, or executing a command.
13 . The method of claim 8 , further comprising determining, based on the time stamp, a wake word duration.
14 . The method of claim 8 , further comprising receiving, from a voice service, one or more responses.
15 . A method comprising:
receiving an audio input comprising a plurality of audio frames;
determining, for each audio frame of the plurality of audio frames, a confidence score that an audio frame of the plurality of audio frames comprises a portion of a wake word;
determining an increase in the confidence score between a first two audio frames of the plurality of audio frames and a decrease in the confidence score between a second two audio frames of the plurality of audio frames;
determining, based on the increase in the confidence score, a first boundary of the wake word, and, based on the decrease in the confidence score, a second boundary of the wake word; and
sending only the audio input between the first boundary and the second boundary to a service.
16 . The method of claim 15 , wherein the audio input comprises one or more user utterances received by a voice enabled device.
17 . The method of claim 16 , further comprising receiving, based on the one or more user utterances, from a cloud based voice service, one or more responses.
18 . The method of claim 15 , further comprising determining, based on the first boundary and the second boundary, a wake word duration.
19 . The method of claim 15 , wherein the service comprises a cloud based voice service.
20 . The method of claim 19 , further comprising determining timing data associated with the first two audio frames and timing data associated with the second two audio frames.