Skip to main navigation Skip to search Skip to main content

Learning Salient Features for Speech Emotion Recognition Using CNN

  • Jiamu Liu
  • , Wenjing Han
  • , Huabin Ruan
  • , Xiaomin Chen
  • , Dongmei Jiang
  • , Haifeng Li
  • Northwestern Polytechnical University Xian
  • Samsung RD Institute China-Beijing (SRC-Beijing)
  • Tsinghua University
  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In this work, a framework based on Convolution Neural Network (CNN) is proposed for speech emotion recognition (SER). We focus on extracting the most salient frames via the proposed CNN structure from the entire frame sequence to represent the utterance. A particular pooling method named global k-max pooling is utilized in our CNN structure (GCNN) to achieve the above objective. We implemented SER experiments on Interactive Emotional Dyadic Motion Capture (IEMOCAP), results are compared to those of some other CNN structures to validate the advancement of the presented framework. The experimental results turn out that GCNN outperforms others CNN models. Besides, experiments are also done to explore how many key frames should be output from GCNN to involve salient emotional information, results illuminate that limited length representation is properer while too long representation is likely containing redundant information decreasing the performance of the model.

Original languageEnglish
Title of host publication2018 1st Asian Conference on Affective Computing and Intelligent Interaction, ACII Asia 2018
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9781538653111
DOIs
StatePublished - 21 Sep 2018
Externally publishedYes
Event1st Asian Conference on Affective Computing and Intelligent Interaction, ACII Asia 2018 - Beijing, China
Duration: 20 May 201822 May 2018

Publication series

Name2018 1st Asian Conference on Affective Computing and Intelligent Interaction, ACII Asia 2018

Conference

Conference1st Asian Conference on Affective Computing and Intelligent Interaction, ACII Asia 2018
Country/TerritoryChina
CityBeijing
Period20/05/1822/05/18

Keywords

  • Convolution neural network
  • Global k-max pooling
  • Representation learning
  • Speech emotion recognition

Fingerprint

Dive into the research topics of 'Learning Salient Features for Speech Emotion Recognition Using CNN'. Together they form a unique fingerprint.

Cite this