System and method for change point detection in multi-media multi-person interactions
One embodiment can provide a method and a system for detecting change points within a conversation. During operation, the system can obtain a signal associated with the conversation and extract a one-dimensional (1D) feature function from the signal. The system can apply Gaussian smoothing on the 1D feature function, identify zero-crossing points on the smoothed 1D feature function, and determine a set of change points within the conversation based on the identified zero-crossing points.
1 . A computer-implemented method for detecting changes in human emotions during a conversation, the method comprising:
recording the conversation using a camera or microphone;
obtaining, by a computer, a signal associated with the recorded conversation;
extracting a one-dimensional (1D) feature function from the signal;
applying Gaussian smoothing on the 1D feature function;
identifying zero-crossing points on the smoothed 1D feature function; and
consolidating the identified zero-crossing points into a smaller set by applying a hierarchical clustering technique on the identified zero-crossing points; and
determining, by the computer, a set of change points indicating the changes in human emotions during the conversation based on the smaller set of zero-crossing points.
2 . The method of claim 1 ,
wherein the signal comprises an audio signal; and
wherein extracting the 1D feature function comprises performing cepstral analysis on the audio signal to obtain one or more Mel-Frequency Cepstral Coefficients (MFCCs).
3 . The method of claim 2 , further comprising:
applying the Gaussian smoothing on a Mel-Frequency Cepstral Coefficient (MFCC);
determining whether a number of identified zero-crossing points on the MFCC is within a predetermined range; and
in response to the number of identified zero-crossing points on the MFCC being outside of the predetermined range, discarding the MFCC and selecting a different MFCC for processing.
4 . The method of claim 2 , further comprising mapping the identified zero-crossing points on the MFCC to time instances.
5 . The method of claim 1 ,
wherein the signal comprises a video signal; and
wherein extracting the 1 D feature function comprises performing facial emotion recognition (FER) analysis on each frame of the video signal to generate a 1D conversational vibe function associated with the video signal.
6 . The method of claim 5 , wherein generating the 1D conversational vibe function further comprises multiplying probability of a detected emotion with a valence value corresponding to the detected emotion.
7 . The method of claim 1 , further comprising annotating the signal using the determined set of change points.
8 . A non-transitory computer-readable storage medium storing instructions that when executed by a processor cause the processor to perform a method for detecting changes in human emotions during a conversation, the method comprising:
configuring a camera or microphone to record the conversation;
obtaining a signal associated with the conversation;
extracting a one-dimensional (1D) feature function from the signal;
applying Gaussian smoothing on the 1D feature function;
identifying zero-crossing points on the smoothed 1D feature function;
consolidating the identified zero-crossing points into a smaller set by applying a hierarchical clustering technique on the identified zero-crossing points; and
determining, by the computer, a set of change points indicating the changes in human emotions during the conversation based on the smaller set of zero-crossing points.
9 . The non-transitory computer-readable storage medium of claim 8 ,
wherein the signal comprises an audio signal; and
wherein extracting the 1 D feature function comprises performing cepstral analysis on the audio signal to obtain one or more Mel-Frequency Cepstral Coefficients (MFCCs).
10 . The non-transitory computer-readable storage medium of claim 9 , wherein the method further comprises:
applying the Gaussian smoothing on a Mel-Frequency Cepstral Coefficient (MFCC);
determining whether a number of identified zero-crossing points on the MFCC is within a predetermined range; and
in response to the number of identified zero-crossing points on the MFCC being outside of the predetermined range, discarding the MFCC and selecting a different MFCC for processing.
11 . The non-transitory computer-readable storage medium of claim 9 , wherein the method further comprises mapping the identified zero-crossing points on the MFCC to time instances.
12 . The non-transitory computer-readable storage medium of claim 8 ,
wherein the signal comprises a video signal; and
wherein extracting the 1D feature function comprises performing facial emotion recognition (FER) analysis on each frame of the video signal to generate a 1D conversational vibe function associated with the video signal.
13 . The non-transitory computer-readable storage medium of claim 12 , wherein generating the 1D conversational vibe function further comprises multiplying probability of a detected emotion with a valence value corresponding to the detected emotion.
14 . The non-transitory computer-readable storage medium of claim 8 ,
wherein the method further comprises annotating the signal using the determined set of change points.
15 . A computer system, comprising:
a processor; and
a storage device storing instructions that when executed by the processor cause the processor to perform a method for detecting changes in human emotions during a conversation, the method comprising:
configuring a camera or microphone to record the conversation;
obtaining a signal associated with the conversation;
extracting a one-dimensional (1D) feature function from the signal;
applying Gaussian smoothing on the 1D feature function;
identifying zero-crossing points on the smoothed 1D feature function;
consolidating the identified zero-crossing points into a smaller set by applying a hierarchical clustering technique on the identified zero-crossing points; and
determining a set of change points indicating the changes in human emotions during the conversation based on the smaller set of zero-crossing points.
16 . The computer system of claim 15 , wherein the method further comprises applying a clustering technique to consolidate the identified zero-crossing points into a smaller set.