TY - GEN
T1 - Multimodal Emotion Recognition in Conversations via Graph Structure Learning
AU - Xiong, Feng
AU - Tu, Geng
AU - Zhang, Yice
AU - Wang, Jun
AU - Chen, Shiwei
AU - Liang, Bin
AU - Yu, Yue
AU - Yang, Min
AU - Xu, Ruifeng
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Multimodal Emotion Recognition in Conversations (MERC) aims to detect emotions expressed in each utterance within conversational videos. Graph-based methods are widely employed in MERC due to their superiority in modeling intricate speaker-sensitive and context-sensitive dependencies in conversations. Despite promising advancements made, existing graph-based methods primarily suffer from two inherent issues due to their reliance on manually predefined graph structures: structural redundancy, which burdens models with irrelevant noise aggregation, and insufficient connections, which results in a lack of cross-modal contextual cues. To address the above issues, we propose a novel graph structure learning framework for MERC, which comprises two key components: Context-aware Graph Sparsification (CGS) and Implicit Graph Relation Mining (IGR). CGS employs an edge selection network to refine the manually predefined graph, filtering out noisy information caused by structural redundancy. IGR explores potential connections that are beneficial for emotional reasoning. Experimental results on two datasets show that our proposed framework significantly improves the performance of graph-based methods in MERC.
AB - Multimodal Emotion Recognition in Conversations (MERC) aims to detect emotions expressed in each utterance within conversational videos. Graph-based methods are widely employed in MERC due to their superiority in modeling intricate speaker-sensitive and context-sensitive dependencies in conversations. Despite promising advancements made, existing graph-based methods primarily suffer from two inherent issues due to their reliance on manually predefined graph structures: structural redundancy, which burdens models with irrelevant noise aggregation, and insufficient connections, which results in a lack of cross-modal contextual cues. To address the above issues, we propose a novel graph structure learning framework for MERC, which comprises two key components: Context-aware Graph Sparsification (CGS) and Implicit Graph Relation Mining (IGR). CGS employs an edge selection network to refine the manually predefined graph, filtering out noisy information caused by structural redundancy. IGR explores potential connections that are beneficial for emotional reasoning. Experimental results on two datasets show that our proposed framework significantly improves the performance of graph-based methods in MERC.
KW - Conversational Understanding
KW - Graph Structure Learning
KW - Multimodal Emotion Recognition
UR - https://www.scopus.com/pages/publications/105022627674
U2 - 10.1109/ICME59968.2025.11209584
DO - 10.1109/ICME59968.2025.11209584
M3 - 会议稿件
AN - SCOPUS:105022627674
T3 - Proceedings - IEEE International Conference on Multimedia and Expo
BT - 2025 IEEE International Conference on Multimedia and Expo
PB - IEEE Computer Society
T2 - 2025 IEEE International Conference on Multimedia and Expo, ICME 2025
Y2 - 30 June 2025 through 4 July 2025
ER -