TY - GEN
T1 - Learning Salient Features for Speech Emotion Recognition Using CNN
AU - Liu, Jiamu
AU - Han, Wenjing
AU - Ruan, Huabin
AU - Chen, Xiaomin
AU - Jiang, Dongmei
AU - Li, Haifeng
N1 - Publisher Copyright:
© 2018 IEEE.
PY - 2018/9/21
Y1 - 2018/9/21
N2 - In this work, a framework based on Convolution Neural Network (CNN) is proposed for speech emotion recognition (SER). We focus on extracting the most salient frames via the proposed CNN structure from the entire frame sequence to represent the utterance. A particular pooling method named global k-max pooling is utilized in our CNN structure (GCNN) to achieve the above objective. We implemented SER experiments on Interactive Emotional Dyadic Motion Capture (IEMOCAP), results are compared to those of some other CNN structures to validate the advancement of the presented framework. The experimental results turn out that GCNN outperforms others CNN models. Besides, experiments are also done to explore how many key frames should be output from GCNN to involve salient emotional information, results illuminate that limited length representation is properer while too long representation is likely containing redundant information decreasing the performance of the model.
AB - In this work, a framework based on Convolution Neural Network (CNN) is proposed for speech emotion recognition (SER). We focus on extracting the most salient frames via the proposed CNN structure from the entire frame sequence to represent the utterance. A particular pooling method named global k-max pooling is utilized in our CNN structure (GCNN) to achieve the above objective. We implemented SER experiments on Interactive Emotional Dyadic Motion Capture (IEMOCAP), results are compared to those of some other CNN structures to validate the advancement of the presented framework. The experimental results turn out that GCNN outperforms others CNN models. Besides, experiments are also done to explore how many key frames should be output from GCNN to involve salient emotional information, results illuminate that limited length representation is properer while too long representation is likely containing redundant information decreasing the performance of the model.
KW - Convolution neural network
KW - Global k-max pooling
KW - Representation learning
KW - Speech emotion recognition
UR - https://www.scopus.com/pages/publications/85055575866
U2 - 10.1109/ACIIAsia.2018.8470393
DO - 10.1109/ACIIAsia.2018.8470393
M3 - 会议稿件
AN - SCOPUS:85055575866
T3 - 2018 1st Asian Conference on Affective Computing and Intelligent Interaction, ACII Asia 2018
BT - 2018 1st Asian Conference on Affective Computing and Intelligent Interaction, ACII Asia 2018
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 1st Asian Conference on Affective Computing and Intelligent Interaction, ACII Asia 2018
Y2 - 20 May 2018 through 22 May 2018
ER -