System and method for identifying sentiment (emotions) in a speech audio input with haptic output
View Patent ↗In a system and method for enabling a user to identify the emotions of speakers to a conversation, spoken audio input is pre-processed using a one-dimensional Mel Spectrogram and/or a two-dimensional Mel-Frequency Cepstral Coefficient (MFCC) matrix, reducing the two-dimensional matrix to a single dimension output, identifying at least one emotion in the audio input using a convolutional or recurrent neural network, and providing the user with haptic feedback corresponding to the at least one emotion in the audio input.
1 . A system for enabling a user to tactilely feel the emotion in a verbal input, said system comprising:
a verbal input receiving device for receiving a spoken input signal;
a processor and memory configured with machine readable code to define a pre-processing stage and an emotion model in the form of a trained artificial neural network stage for extracting at least one emotion associated with the verbal input, and
one or more haptic feedback devices worn by one or more users, which are configured to emit a haptic feedback signal associated with said at least one emotion,
wherein the pre-processing stage is configured to generate a multi-dimensional Mel Spectrogram or Mel-Frequency Cepstral Coefficient (MFCC) matrix from time bands defined in the spoken input signal, and to reduce the multi-dimensional matrix to a single dimensional output by taking the mean value for each time band as well as running mean normalization across all the data and feeding this into the artificial neural network stage, and wherein the haptic feedback signal comprises a vibration that is unique for each emotion or combination of emotions extracted by the neural network stage.
2 . A system of claim 1 , wherein the haptic feedback device comprises a wristband or other wearable item in communication with an output from the artificial neural network stage.
3 . A system of claim 1 , wherein the users are speakers taking part in live, face-to-face conversation, and the verbal input receiving device comprises a microphone.
4 . A system of claim 1 , further comprising representing at least one emotion in auditory form to a speaking participant.
5 . A system of claim 1 , wherein the artificial neural network stage is a recurrent neural network (RNN) that includes layers for performing one or more of:
reducing data overfitting, and transforming data into useful numbers.
6 . A system for enabling a user to tactilely feel the emotion in a verbal input, said system comprising:
a verbal input receiving device for receiving a spoken input signal;
a processor and memory configured with machine readable code to define a pre-processing stage, and an emotion model receiving the output of the pre-processing stage as its input, wherein the emotion model comprises a trained neural network stage for extracting at least one emotion associated with the verbal input,
wherein the pre-processing stage is configured to generate a multi-dimensional Mel Spectrogram or Mel-Frequency Cepstral Coefficient (MFCC) matrix from time bands defined in the spoken input signal, and take the mean value for each time band and also run mean normalization across all the data to reduce the multi-dimensional matrix to a single dimensional output; and
a haptic feedback device worn by a user, which is configured to emit a haptic feedback signal associated with said at least one emotion, wherein the haptic feedback signal comprises a vibration that is unique for each emotion or combination of emotions extracted by the neural network stage.
7 . A system of claim 6 wherein the neural network stage comprises a recurrent neural network (RNN) architecture that includes layers for
recurrently recognizing patterns; and layers to transform data into useful numbers.
8 . A method of analyzing spoken audio input obtained from one or more speaking participants for emotional content, comprising
capturing time bands of the spoken audio input,
pre-processing the time bands of the spoken audio input by transforming the time bands into a multi-dimensional Mel-Frequency Cepstral Coefficient (MFCC) matrix or Mel Spectrogram, and reducing the multi-dimensional matrix to a single-dimension output by taking the mean value for each time band as well as running mean normalization across all the data, the method further comprising,
feeding the single dimensional output into a neural network to identify at least one emotion in the audio input, and
converting the at least one emotion into haptic feedback, wherein the haptic feedback comprises a vibration that is unique for each emotion or combination of emotions identified by the neural network.