TY - GEN
T1 - Multi-modal emotion recognition based on speech and image
AU - Li, Yongqiang
AU - He, Qi
AU - Zhao, Yongping
AU - Yao, Hongxun
N1 - Publisher Copyright:
© Springer International Publishing AG, part of Springer Nature 2018.
PY - 2018
Y1 - 2018
N2 - For the past two decades emotion recognition has gained great attention because of huge potential in many applications. Most works in this field try to recognize emotion from single modal such as image or speech. Recently, there are some studies investigating emotion recognition from multi-modal, i.e., speech and image. The information fusion strategy is a key point for multi-modal emotion recognition, which can be grouped into two main categories: feature level fusion and decision level fusion. This paper explores the emotion recognition from multi-modal, i.e., speech and image. We make a systemic and detailed comparison among several feature level fusion methods and decision level fusion methods such as PCA based feature fusion, LDA based feature fusion, product rule based decision fusion, mean rule based decision fusion and so on. We test all the compared methods on the Surrey Audio-Visual Expressed Emotion (SAVEE) Database. The experimental results demonstrate that emotion recognition based on fusion of speech and image achieved high recognition accuracy than emotion recognition from single modal, and also the decision level fusion methods show superior to feature level fusion methods in this work.
AB - For the past two decades emotion recognition has gained great attention because of huge potential in many applications. Most works in this field try to recognize emotion from single modal such as image or speech. Recently, there are some studies investigating emotion recognition from multi-modal, i.e., speech and image. The information fusion strategy is a key point for multi-modal emotion recognition, which can be grouped into two main categories: feature level fusion and decision level fusion. This paper explores the emotion recognition from multi-modal, i.e., speech and image. We make a systemic and detailed comparison among several feature level fusion methods and decision level fusion methods such as PCA based feature fusion, LDA based feature fusion, product rule based decision fusion, mean rule based decision fusion and so on. We test all the compared methods on the Surrey Audio-Visual Expressed Emotion (SAVEE) Database. The experimental results demonstrate that emotion recognition based on fusion of speech and image achieved high recognition accuracy than emotion recognition from single modal, and also the decision level fusion methods show superior to feature level fusion methods in this work.
KW - Decision level fusion
KW - Feature level fusion
KW - Multi-modal emotion recognition
UR - https://www.scopus.com/pages/publications/85047458691
U2 - 10.1007/978-3-319-77380-3_81
DO - 10.1007/978-3-319-77380-3_81
M3 - 会议稿件
AN - SCOPUS:85047458691
SN - 9783319773797
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 844
EP - 853
BT - Advances in Multimedia Information Processing – PCM 2017 - 18th Pacific-Rim Conference on Multimedia, Revised Selected Papers
A2 - Zeng, Bing
A2 - Li, Hongliang
A2 - El Saddik, Abdulmotaleb
A2 - Fan, Xiaopeng
A2 - Jiang, Shuqiang
A2 - Huang, Qingming
PB - Springer Verlag
T2 - 18th Pacific-Rim Conference on Multimedia, PCM 2017
Y2 - 28 September 2017 through 29 September 2017
ER -