Speech model personalization via ambient context harvesting
An apparatus for speech model with personalization via ambient context harvesting, is described herein. The apparatus includes a microphone, context harvesting module, confidence module, and training module. The context harvesting module is to determine a context associated with the captured audio signals. A confidence module is to determine a confidence of the context as applied to the audio signals. A training module is to train a neural network in response to the confidence being above a predetermined threshold.
1 . An apparatus comprising:
a microphone to capture audio;
interface circuitry;
machine readable instructions stored in a non-transitory memory; and
programmable circuitry to at least one of execute or instantiate the machine readable instructions to:
detect speech based on the audio collected by the microphone;
identify situational data associated with the audio;
determine a context associated with the audio based on the situational data;
determine a confidence score indicative of a likelihood that the speech is associated with the context;
recognize a dialog pattern based on the speech and the situational data, including comparing the speech to reference dialog data, wherein the dialog pattern includes a sequence of interactions in the reference dialog data;
assign a similarity metric to the speech based on the comparison;
classify the speech based on the dialog pattern and the similarity metric; and
based on the confidence score satisfying a confidence threshold, add the speech and the similarity metric to a database.
2 . The apparatus of claim 1 , wherein the situational data includes one or more of a location or a time of day associated with collection of the audio.
3 . The apparatus of claim 1 , wherein the situational data includes the image data representative of an environment in which the audio was collected.
4 . The apparatus of claim 1 , wherein the situational data includes ambient noise in an environment in which the audio was collected.
5 . The apparatus of claim 1 , wherein the programmable circuitry is to:
recognize a speaker associated with the speech; and
identify the situational data based on the speaker.
6 . The apparatus of claim 1 , wherein the database is a database of structured interactions including reference dialog data, and wherein the programmable circuitry is to recognize the dialogue pattern based on a comparison of the speech to the reference dialog data.
7 . The apparatus of claim 6 , wherein the programmable circuitry is to identify the dialog pattern as a sequence of interactions in the reference dialog data.
8 . The apparatus of claim 6 , wherein the programmable circuitry is to assign a similarity metric to the speech based on the comparison.
9 . The apparatus of claim 8 , wherein the programmable circuitry is to include data associated with the classified speech and the similarity metric in training data to train a neural network for speech recognition.
10 . At least one non-transitory memory comprising instructions to cause programmable circuitry to at least:
detect speech from a user based on an audio collected by a microphone;
identify a location associated with collection of the audio;
determine a context associated with the audio based on the location;
determine a confidence score indicative of a likelihood that the speech is associated with the context;
recognize a structured interaction involving the user based on the speech and the location, including comparing the speech to reference dialog data, wherein the structured interaction includes a sequence of interactions in the reference dialog data;
assign a similarity metric to the speech based on the comparison;
classify the speech based on the structured interaction and the similarity metric; and
add the speech and the similarity metric to a database, based on the confidence score satisfying a confidence threshold.
11 . The at least one memory of claim 10 , wherein the instructions cause the programmable circuitry to identify the location based on location data generated by a mobile device.
12 . The at least one memory of claim 10 , wherein the instructions cause the programmable circuitry to identify image data representative of an environment in which the audio was collected.
13 . The at least one memory of claim 10 , wherein the instructions cause the programmable circuitry to:
identify the user associated with the speech; and
recognize the structured interaction based on the identification of the user.
14 . The at least one memory of claim 10 , wherein the instructions cause the programmable circuitry to
access textual data associated with the location, the textual data not associated with the speech; and
generate training data to train a neural network, the training data including the classified speech and the textual data.
15 . An apparatus comprising:
a microphone to capture audio signals;
interface circuitry;
machine readable instructions stored in a non-transitory memory; and
programmable circuitry to at least one of execute or instantiate the machine readable instructions to:
identify a speaker associated with speech;
identify situational data associated with the speech;
determine a context associated with the audio based on the situational data;
determine a confidence score indicative of a likelihood that the speech is associated with the context;
recognize a dialog pattern based on the speech, the identity of the speaker, and the situational data, including comparing the speech to reference dialog data, wherein the dialog pattern includes a sequence of interactions in the reference dialog data; update training data based on the dialog pattern and the confidence score satisfying a confidence threshold, the training data to train a neural network model.
16 . The apparatus of claim 15 , wherein the speaker is a first speaker and the programmable circuitry is to:
recognize the dialogue pattern is an interaction between the first speaker and the second speaker; and
associate at least one of the dialogue pattern or the speech with the interaction.
17 . The apparatus of claim 15 , wherein the situational data includes one or more of a location or a time of day associated with the collection of the speech.
18 . The apparatus of claim 15 , wherein the programmable circuitry is to identify the speaker based on the situational data.
19 . The apparatus of claim 15 , wherein the processor circuitry is to recognize the dialogue pattern based on a comparison of the speech to reference dialog data.
20 . The apparatus of claim 19 , wherein the reference dialog data is associated with the speaker.