Skip to main navigation Skip to search Skip to main content

MMATERIC: Multi-Task Learning and Multi-Fusion for AudioText Emotion Recognition in Conversation

  • Xingwei Liang
  • , You Zou
  • , Xinnan Zhuang
  • , Jie Yang
  • , Taiyu Niu
  • , Ruifeng Xu*
  • *Corresponding author for this work
  • Konka Research Institute
  • Harbin Institute of Technology Shenzhen
  • Northern Arizona University

Research output: Contribution to journalArticlepeer-review

Abstract

The accurate recognition of emotions in conversations helps understand the speaker’s intentions and facilitates various analyses in artificial intelligence, especially in human–computer interaction systems. However, most previous methods need more ability to track the different emotional states of each speaker in a dialogue. To alleviate this dilemma, we propose a new approach, Multi-Task Learning and Multi-Fusion AudioText Emotion Recognition in Conversation (MMATERIC) for emotion recognition in conversation. MMATERIC can refer to and combine the benefits of two distinct tasks: emotion recognition in text and emotion recognition in speech, and production of fused multimodal features to recognize the emotions of different speakers in dialogue. At the core of MATTERIC are three modules: an encoder with multimodal attention, a speaker emotion detection unit (SED-Unit), and a decoder with speaker emotion detection Bi-LSTM (SED-Bi-LSTM). Together, these three modules model the changing emotions of a speaker at a given moment in a conversation. Meanwhile, we adopt multiple fusion strategies in different stages, mainly using model fusion and decision stage fusion to improve the model’s accuracy. Simultaneously, our multimodal framework allows features to interact across modalities and allows potential adaptation flows from one modality to another. Our experimental results on two benchmark datasets show that our proposed method is effective and outperforms the state-of-the-art baseline methods. The performance improvement of our method is mainly attributed to the combination of three core modules of MATTERIC and the different fusion methods we adopt in each stage.

Original languageEnglish
Article number1534
JournalElectronics (Switzerland)
Volume12
Issue number7
DOIs
StatePublished - Apr 2023
Externally publishedYes

Keywords

  • Multi-Task Learning and Multi-Fusion for AudioText Emotion Recognition in Conversation
  • emotion recognition in conversation
  • multi-task learning
  • multimodal fusion

Fingerprint

Dive into the research topics of 'MMATERIC: Multi-Task Learning and Multi-Fusion for AudioText Emotion Recognition in Conversation'. Together they form a unique fingerprint.

Cite this