Bystander speech command rejection for wearable devices
A system comprising a head-wearable apparatus, which includes a plurality of microphones, each microphone of the plurality of microphones being spatially separated from each other microphone of the plurality of microphones, thereby defining a set of spatial separations. The system also includes one or more processors. The system also includes a non-transitory computer readable storage medium including instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including generating, by each microphone of the plurality of microphones, an audio signal corresponding to sound detected by the microphone, thereby generating a plurality of audio signals, and processing the plurality of audio signals, using a trained machine learning model, to generate wearer speech detection data representative of a likelihood that a wearer of the head-wearable apparatus is speaking.
1 . A system comprising:
a head-wearable apparatus, comprising a plurality of microphones, each microphone of the plurality of microphones being spatially separated from each other microphone of the plurality of microphones, thereby defining a set of spatial separations;
one or more processors; and
a non-transitory computer readable storage medium comprising instructions that when executed by the one or more processors cause the one or more processors to perform operations comprising:
generating, by each individual microphone of the plurality of microphones, an audio signal corresponding to sound detected by the individual microphone, thereby generating a plurality of audio signals; and
processing the plurality of audio signals, using a trained machine learning model, to generate wearer speech detection data representative of a likelihood that a wearer of the head-wearable apparatus is speaking;
wherein the trained machine learning model is trained using supervised learning based on a training dataset comprising a plurality of labelled multi-channel audio samples generated by:
generating a plurality of directional sounds, each directional sound originating from a different direction relative to a plurality of spatially separated microphones, the plurality of directional sounds including:
one or more wearer sounds originating from a direction corresponding to a wearer speech location; and
a plurality of bystander sounds originating from a direction corresponding to a location other than the wearer speech location; and
for each individual directional sound of the plurality of directional sounds:
for each individual recording microphone of the plurality of spatially separated recording microphones:
generating an audio recording, by the individual recording microphone, of the individual directional sound, thereby generating a plurality of audio recordings;
processing the plurality of audio recordings to generate a plurality of DIR filters corresponding to the plurality of directional sounds;
applying the plurality of DIR filters to a plurality of source audio samples to generate a plurality of multi-channel audio samples, each multi-channel audio sample comprising an audio channel for each microphone of the plurality of spatially separated microphones; and
associating each of a plurality of the multi-channel audio samples with a respective label representative of whether a wearer sound corresponds to an individual DIR filter of the plurality of DIR filters, thereby generating the plurality of labelled multi-channel audio samples.
2 . The system of claim 1 , wherein two microphones of the plurality of microphones are located relative to the head-wearable apparatus such that a spatial vector passing through both of the two microphones passes through a wearer speech location where sound originates when the wearer is speaking.
3 . The system of claim 2 , wherein:
the head-wearable apparatus comprises a pair of augmented reality glasses having:
a pair of frames; and
a pair of stems extending backward from the pair of frames when worn; and
the two microphones of the plurality of microphones are located on one of the pair of stems and a lower portion of one of the pair of frames, respectively.
4 . The system of claim 3 , wherein:
the plurality of microphones comprises:
a first microphone in a first location on a left stem of the pair of stems;
a second microphone in a second location on the left stem, behind the first location;
a third microphone in a third location on a lower portion of a left frame of the pair of frames;
a fourth microphone in a fourth location on a right stem of the pair of stems;
a fifth microphone in a fifth location on the right stem, behind the fourth location; and
a sixth microphone in a sixth location on a lower portion of a right frame of the pair of frames.
5 . The system of claim 1 , wherein the plurality of spatially separated recording microphones are spatially separated from each other by a same set of spatial separations as the plurality of microphones of the head-wearable apparatus.
6 . The system of claim 1 , wherein generating the plurality of labelled multi-channel audio samples further comprises:
combining one or more noise profiles with the plurality of labelled multi-channel audio samples, at one or more signal to noise ratios, to generate noisy multi-channel audio samples, each noisy multi-channel audio sample being an audio samples at one of the one or more signal to noise ratios and having one of the one or more noise profiles.
7 . The system of claim 1 , wherein generating the plurality of labelled multi-channel audio samples further comprises:
excluding from the training dataset one or more multi-channel audio samples generated using a DIR filter corresponding to a bystander sound.
8 . The system of claim 1 , wherein the trained machine learning model is a classifier comprising a convolutional neural network (CNN).
9 . The system of claim 8 , wherein the trained machine learning model is further trained by:
generating a plurality of spectrograms based on the plurality of labelled multi-channel audio samples;
using the plurality of spectrograms as training inputs; and
using labels of the plurality of labelled multi-channel audio samples as training labels.
10 . A method comprising:
training a machine learning model to generate a trained machine learning model using supervised learning based on a training dataset comprising a plurality of labelled multi-channel audio samples generated by:
generating a plurality of directional sounds, each individual directional sound originating from a different direction relative to a plurality of spatially separated recording microphones, the plurality of directional sounds including:
one or more wearer sounds originating from a direction corresponding to a wearer speech location; and
a plurality of bystander sounds originating from a direction corresponding to a location other than the wearer speech location; and
for each individual directional sound of the plurality of directional sounds:
for each individual recording microphone of the plurality of spatially separated recording microphones:
generating an audio recording, by the individual recording microphone, of the individual directional sound, thereby generating a plurality of audio recordings;
processing the plurality of audio recordings to generate a plurality of DIR filters corresponding to the plurality of directional sounds;
applying the plurality of DIR filters to a plurality of source audio samples to generate a plurality of multi-channel audio samples, each multi-channel audio sample comprising an audio channel for each microphone of the plurality of spatially separated recording microphones; and
associating each of the plurality of the multi-channel audio samples with a respective label representative of whether a wearer sound corresponds to an individual DIR filter of the plurality of DIR filters, thereby generating the plurality of labelled multi-channel audio samples;
generating, by each individual microphone of a plurality of microphones of a head-wearable apparatus, an audio signal corresponding to sound detected by the individual microphone, thereby generating a plurality of audio signals;
the individual microphone of the plurality of microphones being spatially separated from each other microphone of the plurality of microphones, thereby defining a set of spatial separations; and
processing the plurality of audio signals, using the trained machine learning model, to generate wearer speech detection data representative of a likelihood that a wearer of the head-wearable apparatus is speaking.
11 . The method of claim 10 , wherein two microphones of the plurality of microphones are located relative to the head-wearable apparatus such that a spatial vector passing through both of the two microphones passes through a wearer speech location where sound originates when the wearer is speaking.
12 . The method of claim 11 , wherein:
the head-wearable apparatus comprises a pair of augmented reality glasses having:
a pair of frames; and
a pair of stems extending backward from the pair of frames when worn; and
the two microphones of the plurality of microphones are located on one of the pair of stems and a lower portion of one of the pair of frames, respectively.
13 . The method of claim 10 , wherein the plurality of spatially separated recording microphones are spatially separated from each other by a same set of spatial separations as the plurality of microphones of the head-wearable apparatus.
14 . The method of claim 10 , wherein generating the plurality of labelled multi-channel audio samples further comprises:
combining one or more noise profiles with the plurality of labelled multi-channel audio samples, at one or more signal to noise ratios, to generate noisy multi-channel audio samples, each noisy multi-channel audio sample being an audio samples at one of the one or more signal to noise ratios and having one of the one or more noise profiles.
15 . The method of claim 10 , wherein generating the plurality of labelled multi-channel audio samples further comprises:
excluding from the training dataset one or more multi-channel audio samples generated using a DIR filter corresponding to a bystander sound.
16 . The method of claim 10 , wherein the trained machine learning model is a classifier comprising a convolutional neural network (CNN).
17 . The method of claim 16 , wherein the trained machine learning model is further trained by:
generating a plurality of spectrograms based on the plurality of labelled multi-channel audio samples;
using the plurality of spectrograms as training inputs; and
using labels of the plurality of labelled multi-channel audio samples as training labels.
18 . A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a system, cause the system to:
train a machine learning model to generate a trained machine learning model using supervised learning based on a training dataset comprising a plurality of labelled multi-channel audio samples generated by:
generating a plurality of directional sounds, each individual directional sound originating from a different direction relative to a plurality of spatially separated recording microphones, the plurality of directional sounds including:
one or more wearer sounds originating from a direction corresponding to a wearer speech location; and
a plurality of bystander sounds originating from a direction corresponding to a location other than the wearer speech location; and
for each individual directional sound of the plurality of directional sounds:
for each individual recording microphone of the plurality of spatially separated recording microphones:
generating an audio recording, by the individual recording microphone, of the individual directional sound, thereby generating a plurality of audio recordings;
processing the plurality of audio recordings to generate a plurality of DIR filters corresponding to the plurality of directional sounds;
applying the plurality of DIR filters to a plurality of source audio samples to generate a plurality of multi-channel audio samples, each multi-channel audio sample comprising an audio channel for each microphone of the plurality of spatially separated recording microphones; and
associating each of a plurality of the multi-channel audio samples with a respective label representative of whether a wearer sound corresponds to an individual DIR filter of the plurality of DIR filters, thereby generating the plurality of labelled multi-channel audio samples;
generate, by each individual microphone of a plurality of microphones of a head-wearable apparatus, an audio signal corresponding to sound detected by the individual microphone, thereby generating a plurality of audio signals;
the individual microphone of the plurality of microphones being spatially separated from each other microphone of the plurality of microphones, thereby define a set of spatial separations; and
process the plurality of audio signals, using a trained machine learning model, to generate wearer speech detection data representative of a likelihood that a wearer of the head-wearable apparatus is speaking.
19 . The non-transitory computer-readable storage medium of claim 18 , wherein two microphones of the plurality of microphones are located relative to the head-wearable apparatus such that a spatial vector passing through both of the two microphones passes through the wearer speech location where sound originates when the wearer is speaking.
20 . The non-transitory computer-readable storage medium of claim 19 , wherein:
the head-wearable apparatus comprises a pair of augmented reality glasses having:
a pair of frames; and
a pair of stems extending backward from the pair of frames when worn; and
the two microphones of the plurality of microphones are located on one of the pair of stems and a lower portion of one of the pair of frames, respectively.