IP Library › Granted Patent US 11,837,252
Granted Patent B2
US 11,837,252 · App. 17/845,908 · Granted Dec 5, 2023

Speech emotion recognition method and system based on fused population information

Inventors: Taihao Li (Zhejiang, CN); Shukai Zheng (Zhejiang, CN); Yulong Liu (Zhejiang, CN); Guanxiong Pei (Zhejiang, CN); Shijie Ma (Zhejiang, CN)
Assignee: Zhejiang Lab
G10L25/63G10L25/18G10L25/21G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,837,252
App. No.
17/845,908
Granted
Dec 5, 2023
Kind
B2
Abstract

The present invention discloses a speech emotion recognition method and system based on fused population information. The method includes the following steps: S 1 : acquiring a user's audio data; S 2 : preprocessing the audio data, and obtaining a Mel spectrogram feature; S 3 : cutting off a front mute segment and a rear mute segment of the Mel spectrogram feature; S 4 : obtaining population depth feature information through a population classification network; S 5 : obtaining Mel spectrogram depth feature information through a Mel spectrogram preprocessing network; S 6 : fusing the population depth feature information and the Mel spectrogram depth feature information through SENet to obtain fused information; and S 7 : obtaining an emotion recognition result from the fused information through a classification network.

Claims (91)

1. A speech emotion recognition method based on fused population information, comprising the following steps:

S 1 : acquiring a user's audio data, expressed as X audio , through a recording acquisition device;

S 2 : preprocessing the acquired audio data X audio to generate a Mel spectrogram feature, expressed as X mel ;

S 3 : calculating energy of Mel spectrograms in different time frames for the generated Mel spectrogram feature X mel , cutting off a front mute segment and a rear mute segment by setting a threshold to obtain a Mel spectrogram feature, expressed as X input , with a length of T;

S 4 : inputting the Mel spectrogram feature X input obtained in S 3 into a population classification network to obtain population depth feature information, expressed as H p ;

S 5 : inputting the Mel spectrogram feature X input obtained in S 3 into a Mel spectrogram preprocessing network to obtain Mel spectrogram depth feature information, expressed as H m ;

S 6 : fusing the population depth feature information H p extracted in S 4 with the Mel spectrogram depth feature information H m extracted in S 5 through a channel attention network SENet to obtain a fused feature, expressed as H f ; and

S 7 : inputting the fused feature H f in S 6 into the population classification network through a pooling layer to perform emotion recognition;

the population classification network is composed of a three-layer Long Short Term Memory (LSTM) network structure, and the S 4 specifically comprises the following steps:

S 4 _ 1 : first, segmenting the inputted Mel spectrogram feature X input with the length of T into three Mel spectrogram segments

T

2

in equal length in an overlapped manner, wherein the segmentation method is as follows: 0 to

T

2

is segmented as a first segment,

T

4

to

3

⁢

T

4

is segmented as a second segment, and

T

2

to T is segmented as a third segment; and

S 4 _ 2 : inputting the three Mel spectrogram segments segmented in S 4 _ 1 into the three-layer LSTM network in turn, then taking the last output from the three-layer LSTM network as a final state, obtaining three hidden features for the three Mel spectrogram segments at last, and finally averaging the three hidden features to obtain the final population feature information H p .

2. The speech emotion recognition method based on fused population information of claim 1 , wherein the Mel spectrogram preprocessing network in the S 5 is composed of a ResNet network and a feature map scaling (FMS) network which are cascaded, and the S 5 specifically comprises the following steps:

first, expanding the Mel spectrogram feature X input with the length of T into a 3D matrix;

second, extracting emotion-related information from the Mel spectrogram feature X input by using the ResNet network structure and adopting a two-layer convolution and maximum pooling structure; and

third, effectively combining the emotion-related information extracted by the ResNet network through an FMS network architecture to finally obtain the Mel spectrogram depth feature information H m .

3. The speech emotion recognition method based on fused population information of claim 1 , wherein the S 6 specifically comprises the following steps:

S 6 _ 1 : the population depth feature information H p is a 1D vector in space R C , where C represents a channel dimension; the Mel spectrogram depth feature information H m is a 3D matrix in space R T×W×C , where T represents a time dimension, W represents a width dimension, and C represents the channel dimension; performing global average pooling on the Mel spectrogram depth feature information H m in the time dimension T and the width dimension W through the SENet network, and converting the Mel spectrogram depth feature information H m into a C-dimensional vector to obtain a 1D vector H p_avg in the space R C ; wherein

H m =[H 1 ,H 2 ,H 3 , . . . ,H C ]

where,

H c =└[h 1,1 c ,h 2,1 c ,h 3,1 c . . . ,h T,1 c ] T ,[h 1,2 c ,h 2,2 c ,h 3,2 c . . . ,h T,2 c ] T , . . . ,[h 1,W c ,h 2,W c ,h 3,W c . . . ,h T,W c ] T ┘

in addition,

H p_avg =[h p_avg 1 ,h p_avg 2 ,h p_avg 3 , . . . ,h p_avg C ]

a formula of the global average pooling is as follows:

h

p

⁢

_

⁢

avq

c

=

1

TW

⁢

∑

i

=

1

,

j

=

1

T

,

W

h

ij

c

S 6 _ 2 : splicing the 1D vector H p_avg obtained in S 6 _ 1 with the population depth feature information H p to obtain a spliced feature, expressed as H c , wherein

H c =└H p_avg ,H p ┘

S 6 _ 3 : inputting the spliced feature H c obtained in S 6 _ 2 into a two-layer fully-connected network to obtain a channel weight vector W c , where a calculation formula of the two-layer fully-connected network is as follows:

Y=W*X+b

where Y represents an output of the two-layer fully-connected network, X represents an input of the two-layer fully-connected network, W represents a weighting parameter of the two-layer fully-connected network, and b represents a bias parameter of the two-layer fully-connected network; and

S 6 _ 4 : multiplying the channel weight vector W c obtained in S 6 _ 3 by the Mel spectrogram depth feature information H m obtained in S 5 to obtain an emotion feature matrix, and performing global average pooling on the emotion feature matrix in a dimension of T×W to obtain a fused feature, expressed as H f .

4. The speech emotion recognition method based on fused population information of claim 1 , wherein the S 7 specifically comprises the following steps:

S 7 _ 1 : after passing through the pooling layer, inputting the fused feature H f obtained in S 6 into the two-layer fully-connected network to obtain a 7-dimensional feature vector H b , where 7 represents a number of all emotion categories; and

S 7 _ 2 : taking the 7-dimensional feature vector H b =[h b 1 , h b 2 , h b 3 , h b 4 , h b 5 , h b 6 , h b 7 ] obtained in S 7 _ 1 as an independent variable of a Softmax operator, calculating a final value of Softmax as a probability value of an inputted audio belonging to each emotion category, and finally selecting the category with the maximum probability value as a final audio emotion category, wherein a calculation formula of Softmax is as follows:

p

i

=

e

h

b

i

∑

7

n

=

1

e

h

b

n

where e is a constant.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2022
From: LI, TAIHAO; ZHENG, SHUKAI; LIU, YULONG; PEI, GUANXIONG; MA, SHIJIE
To: ZHEJIANG LAB
Reel/Frame 060361/0372 →
Priority Claims (1)
CN 202110322720.X · Mar 26, 2021 · national
Continuity (2)
Continuation PCTCN2022070728 · Jan 7, 2022
Related Publication 20220328065A1 · Oct 13, 2022