System and method for identifying sentiment (emotions) in a speech audio input
View Patent ↗In a system and method for enabling a user to identify the emotions of speakers during a telephone or online conversation, spoken audio input is pre-processed using a one-dimensional Mel Spectrogram and/or a two-dimensional Mel-Frequency Cepstral Coefficient (MFCC) matrix, reducing the two-dimensional matrix to a single dimension output, and identifying at least one emotion in the audio input using a convolutional or recurrent neural network.
1 . A method of analyzing spoken audio input for emotional content, comprising
capturing a time series of the spoken audio input and defining time bands in the audio input,
pre-processing the time series of the spoken audio input to generate a multi-dimensional Mel Spectrogram, or Mel-Frequency Cepstral Coefficient (MFCC) matrix,
reducing the multi-dimensionality of the Mel Spectrogram or MFCC matrix to a single-dimension matrix output by taking the mean value for each time band as well as running mean normalization across all the data to define the single-dimension matrix output,
feeding the single-dimension matrix output from the pre-processing into a trained neural network, which provides an output that identifies at least one emotion in the audio input, and
presenting the emotion visually on a display screen in the form of a written description or graphic depiction.
2 . A method of claim 1 , wherein the audio input comprises an analog input or a digital input configured to define a pre-defined number of frequency values.
3 . A method of claim 1 , further comprising representing the emotional content of each speaker in auditory form.
4 . A method of claim 1 , wherein the artificial neural network comprises a recurrent neural network (RNN) that includes layers for performing one or more of the following steps:
transforming data into useful numbers, and reducing data overfitting.
5 . A system for analyzing emotions in speech, comprising
a verbal input receiving device for receiving a spoken input signal,
a processor and memory configured with machine readable code to define a pre-processing stage, wherein the pre-processing stage is configured to generate a multi-dimensional Mel Spectrogram or Mel-Frequency Cepstral Coefficient (MFCC) matrix from time bands defined in the spoken input signal, and wherein the pre-processing stage takes the mean value for each time band as well as running mean normalization across all the data to define a single dimensional output, the system further comprising an emotion model in the form of a trained multi-layered neural network arranged to receive as its input, the single dimensional output from the pre-processing stage, wherein the neural network is configured either as a convolutional neural network or as a recurrent neural network, and provides as its output, data defining one more emotions in the speech.
6 . A system of claim 5 , wherein the neural network comprises a recurrent neural network with layers to transform data into useful numbers, and layers to reduce data overfitting.
7 . A system for analyzing spoken audio input from one or more participants, for emotional content, comprising
a pre-processing stage configured to generate a multi-dimensional Mel Spectrogram or-Mel-Frequency Cepstral Coefficient (MFCC) matrix from time bands defined in the spoken audio input, wherein the pre-processing stage takes the mean value for each time band as well as running mean normalization across all the data to define a single-dimensional output,
a trained convolutional neural network (CNN) or a trained recurrent neural network (RNN), arranged to receive as its input the single-dimensional output from the pre-processing stage, and configured to identify at its output at least one emotion in the audio input, the system further comprising
one or more display screens for representing the at least one emotion from the output of the CNN or RNN in visual form to one or more participants by means of one or more of written description, and graphic representation.
8 . A system of claim 7 , wherein the participants are speakers taking part in a telephone or online conversation.
9 . A system of claim 7 , further comprising an audio output for representing the at least one emotion in auditory form to a participant.
10 . A system of claim 7 , wherein the system is part of an online conference call network and the participants are connected to the network by means of user access devices.
11 . A system of claim 10 , wherein the user access devices include one or more of cell phone, tablet, laptop, or desktop computer.
12 . A method of analyzing spoken audio input obtained from one or more speaking participants for emotional content, comprising
capturing time bands of the spoken audio input,
processing the time bands of the spoken audio input by transforming the time bands into a multi-dimensional Mel Spectrogram or Mel-Frequency Cepstral Coefficient (MFCC) matrix, and reducing the multi-dimensional matrix from the Mel Spectrogram or MFCC matrix to a single-dimensional matrix output by taking the mean value for each time band as well as running mean normalization across all the data to define the single-dimensional output, the method further comprising,
feeding the single-dimensional output into a trained neural network, wherein the trained neural network comprises either a convolutional neural network (CNN) or a recurrent neural network (RNN), wherein the output from the CNN or RNN defines one or more emotions in the spoken audio input, and
presenting the one or more emotions visually on one or more display screens by means of one or more of written description, and graphic representation, or by presenting the emotions of one or more of the participants by way of verbal feedback to one or more of the participants.
13 . A system of claim 12 , wherein, in the case of a Recurrent Neural Network architecture (RNN), the RNN includes one or more of:
layers to transform data into useful numbers, and
layers to reduce data overfitting.