Skip to main navigation Skip to search Skip to main content

Multi-modal emotion recognition based on speech and image

  • School of Electrical Engineering and Automation, Harbin Institute of Technology
  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

For the past two decades emotion recognition has gained great attention because of huge potential in many applications. Most works in this field try to recognize emotion from single modal such as image or speech. Recently, there are some studies investigating emotion recognition from multi-modal, i.e., speech and image. The information fusion strategy is a key point for multi-modal emotion recognition, which can be grouped into two main categories: feature level fusion and decision level fusion. This paper explores the emotion recognition from multi-modal, i.e., speech and image. We make a systemic and detailed comparison among several feature level fusion methods and decision level fusion methods such as PCA based feature fusion, LDA based feature fusion, product rule based decision fusion, mean rule based decision fusion and so on. We test all the compared methods on the Surrey Audio-Visual Expressed Emotion (SAVEE) Database. The experimental results demonstrate that emotion recognition based on fusion of speech and image achieved high recognition accuracy than emotion recognition from single modal, and also the decision level fusion methods show superior to feature level fusion methods in this work.

Original languageEnglish
Title of host publicationAdvances in Multimedia Information Processing – PCM 2017 - 18th Pacific-Rim Conference on Multimedia, Revised Selected Papers
EditorsBing Zeng, Hongliang Li, Abdulmotaleb El Saddik, Xiaopeng Fan, Shuqiang Jiang, Qingming Huang
PublisherSpringer Verlag
Pages844-853
Number of pages10
ISBN (Print)9783319773797
DOIs
StatePublished - 2018
Externally publishedYes
Event18th Pacific-Rim Conference on Multimedia, PCM 2017 - Harbin, China
Duration: 28 Sep 201729 Sep 2017

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume10735 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference18th Pacific-Rim Conference on Multimedia, PCM 2017
Country/TerritoryChina
CityHarbin
Period28/09/1729/09/17

Keywords

  • Decision level fusion
  • Feature level fusion
  • Multi-modal emotion recognition

Fingerprint

Dive into the research topics of 'Multi-modal emotion recognition based on speech and image'. Together they form a unique fingerprint.

Cite this