TY - GEN
T1 - A Graph-Enhanced MLLM for Hierarchical Multimodal Emotion Understanding and Support in Conversations
AU - Tu, Geng
AU - Niu, Taiyu
AU - Zeng, Xi
AU - Xu, Ruifeng
AU - Zhang, Min
N1 - Publisher Copyright:
© 2026 Owner/Author.
PY - 2026/7/19
Y1 - 2026/7/19
N2 - Multimodal emotion dialogue system involves a hierarchical progression: Perception (Multimodal Emotion Recognition in Conversation, MERC), Reasoning (Multimodal Emotion-Cause Pair Extraction, MECPE), and Support (Multimodal Emotional Support Conversation, MESC). Although existing methods have shifted toward MLLM-based paradigms with stronger generalization ability, they still handle MERC, MECPE, and MESC as independent components. This decoupled design disrupts the hierarchical relationship among perception, reasoning, and support. Moreover, their multimodal fusion relies on coarse-grained concatenation of encoder features into the prompt, ignoring structured relational dependencies among speakers and modalities. To address these limitations, we propose GraphEMO, a unified multimodal emotion understanding and support framework based on a graph-enhanced MLLM. It models MERC, MECPE, and MESC as a residual chain of task-specific priors, enabling MESC to build on perceptual and causal representations from preceding stages. In addition, we leverages dual graph attention to capture cross- and intra-modal relational dynamics among speakers, enhancing structured relational modeling. Experiments on MERC, MECPE, and MESC benchmarks show that GraphEMO consistently improves performance on Qwen and LLaMA models, achieving state-of-the-art results.
AB - Multimodal emotion dialogue system involves a hierarchical progression: Perception (Multimodal Emotion Recognition in Conversation, MERC), Reasoning (Multimodal Emotion-Cause Pair Extraction, MECPE), and Support (Multimodal Emotional Support Conversation, MESC). Although existing methods have shifted toward MLLM-based paradigms with stronger generalization ability, they still handle MERC, MECPE, and MESC as independent components. This decoupled design disrupts the hierarchical relationship among perception, reasoning, and support. Moreover, their multimodal fusion relies on coarse-grained concatenation of encoder features into the prompt, ignoring structured relational dependencies among speakers and modalities. To address these limitations, we propose GraphEMO, a unified multimodal emotion understanding and support framework based on a graph-enhanced MLLM. It models MERC, MECPE, and MESC as a residual chain of task-specific priors, enabling MESC to build on perceptual and causal representations from preceding stages. In addition, we leverages dual graph attention to capture cross- and intra-modal relational dynamics among speakers, enhancing structured relational modeling. Experiments on MERC, MECPE, and MESC benchmarks show that GraphEMO consistently improves performance on Qwen and LLaMA models, achieving state-of-the-art results.
KW - conversational emotion recognition
KW - emotion-cause pair extraction
KW - emotional support conversation
UR - https://www.scopus.com/pages/publications/105047237663
U2 - 10.1145/3805712.3809910
DO - 10.1145/3805712.3809910
M3 - 会议稿件
AN - SCOPUS:105047237663
T3 - SIGIR 2026 - Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval
SP - 4199
EP - 4204
BT - SIGIR 2026 - Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval
PB - Association for Computing Machinery, Inc
T2 - 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2026
Y2 - 20 July 2026 through 24 July 2026
ER -