Systems, methods, and devices for low-power audio signal detection
Systems, methods, and devices detect wake signals included in audio signals. Methods include receiving a dataset including raw audio data, the raw audio data comprising a plurality of audio samples and associated metadata, and generating, using one or more processing elements, an augmented dataset based on the raw audio data, the augmented dataset comprising a plurality of annotations identifying types of raw audio data. Methods further include generating, using the one or more processing elements, a feature dataset by extracting features from the augmented dataset based, at least in part, on the plurality of annotations, and generating, using the one or more processing elements, a wake signal detection model based, at least in part, on the feature dataset, the wake signal detection model being a machine learning model trained based on the feature dataset.
1 . A method comprising:
receiving a dataset including raw audio data, the raw audio data comprising a plurality of audio samples and associated metadata;
generating, using one or more processing elements, an augmented dataset based on the raw audio data, the augmented dataset comprising a plurality of annotations identifying types of raw audio data, the plurality of annotations being generated based, at least in part, on the metadata associated with the plurality of audio samples;
generating, using the one or more processing elements, a feature dataset by extracting features from the augmented dataset based, at least in part, on the plurality of annotations and using the plurality of annotations to serialize the extracted features to generate an input for a wake signal detection model; and
generating, using the one or more processing elements, the wake signal detection model based, at least in part, on the feature dataset, the wake signal detection model being a machine learning model trained based on the feature dataset, the generating of the wake signal detection model further comprising reducing a number of dimensions of the machine learning model based, at least in part, on power consumption characteristics of a target audio signal processing device.
2 . The method of claim 1 further comprising:
generating, using the one or more processing elements, training data based, at least in part, on the feature dataset.
3 . The method of claim 2 , wherein the generating of the training data further comprises:
concatenating at least some of the extracted features included in the feature dataset; and
generating an output file based on the concatenation of the at least some of the extracted features.
4 . The method of claim 1 , wherein the generating of the augmented dataset further comprises:
classifying a plurality of phonemes included in the raw audio data;
generating a plurality of tokens based on the plurality of phonemes; and
generating the plurality of annotations based on the plurality of tokens.
5 . The method of claim 4 , wherein the classifying is performed by an automatic speech recognition model, and wherein the plurality of tokens identify whether or not speech is present in each of the plurality of phonemes.
6 . The method of claim 1 further comprising:
testing the wake signal detection model using test data.
7 . The method of claim 6 further comprising:
modifying one or more weights associated with the wake signal detection model based on a result of the testing.
8 . The method of claim 1 further comprising:
generating a low-power model based on the wake signal detection model, the low-power model having the reduced number of dimensions.
9 . The method of claim 8 , wherein the low-power model is configured to execute on a low-power device in real-time.
10 . A system comprising:
a communications interface configured to receive raw audio data, the raw audio data comprising a plurality of audio samples and associated metadata;
one or more processing elements configured to:
generate a dataset based on the raw audio data received from the communications interface;
generate an augmented dataset based on the raw audio data, the augmented dataset comprising a plurality of annotations identifying types of raw audio data, the plurality of annotations being generated based, at least in part, on the metadata associated with the plurality of audio samples;
generate a feature dataset by extracting features from the augmented dataset based, at least in part, on the plurality of annotations and using the plurality of annotations to serialize the extracted features to generate an input for a wake signal detection model; and
generate the wake signal detection model based, at least in part, on the feature dataset, the wake signal detection model being a machine learning model trained based on the feature dataset, the generating of the wake signal detection model further comprising reducing a number of dimensions of the machine learning model based, at least in part, on power consumption characteristics of a target audio signal processing device.
11 . The system of claim 10 , wherein the one or more processing elements are further configured to:
generate training data based, at least in part, on the feature dataset.
12 . The system of claim 11 , wherein the one or more processing elements are further configured to:
concatenate at least some of the extracted features included in the feature dataset; and
generate an output file based on the concatenation of the at least some of the extracted features.
13 . The system of claim 10 , wherein the one or more processing elements are further configured to:
classify a plurality of phonemes included in the raw audio data, wherein the classifying is performed by an automatic speech recognition model;
generate a plurality of tokens based on the plurality of phonemes, wherein the plurality of tokens identify whether or not speech is present in each of the plurality of phonemes; and
generate the plurality of annotations based on the plurality of tokens.
14 . The system of claim 10 , wherein the one or more processing elements are further configured to:
generate a low-power model based on the wake signal detection model, the low-power model having the reduced number of dimensions.
15 . The system of claim 14 , wherein the low-power model is configured to execute on a low-power device in real-time.
16 . A device comprising:
one or more processing elements configured to:
receive a dataset including raw audio data, the raw audio data comprising a plurality of audio samples and associated metadata;
generate an augmented dataset based on the raw audio data, the augmented dataset comprising a plurality of annotations identifying types of raw audio data, the plurality of annotations being generated based, at least in part, on the metadata associated with the plurality of audio samples;
generate a feature dataset by extracting features from the augmented dataset based, at least in part, on the plurality of annotations and using the plurality of annotations to serialize the extracted features to generate an input for a wake signal detection model; and
generate the wake signal detection model based, at least in part, on the feature dataset, the wake signal detection model being a machine learning model trained based on the feature dataset, the generating of the wake signal detection model further comprising reducing a number of dimensions of the machine learning model based, at least in part, on power consumption characteristics of a target audio signal processing device.
17 . The device of claim 16 , wherein the one or more processing elements are further configured to:
generate training data based, at least in part, on the feature dataset.
18 . The device of claim 17 , wherein the one or more processing elements are further configured to:
concatenate at least some of the extracted features included in the feature dataset; and
generate an output file based on the concatenation of the at least some of the extracted features.
19 . The device of claim 16 , wherein the one or more processing elements are further configured to:
classify a plurality of phonemes included in the raw audio data, wherein the classifying is performed by an automatic speech recognition model;
generate a plurality of tokens based on the plurality of phonemes, wherein the plurality of tokens identify whether or not speech is present in each of the plurality of phonemes; and
generate the plurality of annotations based on the plurality of tokens.
20 . The device of claim 16 , wherein the one or more processing elements are further configured to:
generate a low-power model based on the wake signal detection model, wherein the low-power model has the reduced number of dimensions, and wherein the low-power model is configured to execute on a low-power device in real-time.