Skip to main navigation Skip to search Skip to main content

Is multimodal conversational emotion recognition satisfactory? Exploring the gaps in performance, generalization, and confidence

  • Geng Tu
  • , Ran Jing
  • , Xuan Luo
  • , Erik Cambria
  • , Wenjie Li
  • , Ruifeng Xu*
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen
  • Peng Cheng Laboratory
  • Shenzhen Loop Area Institute
  • Nanyang Technological University
  • Hong Kong Polytechnic University

Research output: Contribution to journalArticlepeer-review

Abstract

Recent advances in Multimodal Emotion Recognition in Conversations (MERC) focus on speaker-aware context modeling and multimodal fusion. However, there remain three challenges: (1) Lack of standardized evaluation protocols complicates fair model comparison; (2) Static datasets fixed speakers, topics, and corpora hinder generalization to unseen scenarios; (3) Overconfident predictions undermine reliability. To address these, this paper explores three critical questions: (1) Are existing MERC models performing adequately? (2) Can they generalize to diverse scenarios? (3) Can we trust their confidence? Through a rigorous reassessment of existing models, including tests on unseen scenarios, we identify key strengths, weaknesses, and synergies across different models, while highlighting the crucial role of confidence calibration in improving model reliability and fairness.2

Original languageEnglish
Article number113087
JournalPattern Recognition
Volume175
DOIs
StatePublished - Jul 2026
Externally publishedYes

Keywords

  • Calibration
  • Conversational emotion recognition
  • Generalization

Fingerprint

Dive into the research topics of 'Is multimodal conversational emotion recognition satisfactory? Exploring the gaps in performance, generalization, and confidence'. Together they form a unique fingerprint.

Cite this