Abstract
Incomplete multimodal learning has become a critical challenge in Emotion Recognition in Conversation (ERC). However, most methods focus on individual utterances instead of conversational data, thus often overlooking the rich contextual information embedded within conversations. Moreover, existing incomplete ERC approaches are typically limited to joint multimodal representation learning for missing modality recovery, failing to account for the complementary information inherent in modality-specific characteristics. In this paper, we propose a Context-guided Recovering Network (CRNet) to comprehensively explore both modality characteristics and contextual information in conversational data for incomplete multimodal ERC. Specifically, to effectively leverage cross-modal semantic consistency while preserving the modality-specific information, we design a context consistent recovery module. This module uses the semantic relations between utterances to guide the recovery of missing components. Furthermore, to maintain the balance between preserving emotion consistency and uncovering emotion shifts in contextual modeling, we introduce a dual-level semantic learning module, which is composed of the instance-level contrastive unit and label-based emotion shift unit. The instance-level contrastive unit enforces similarity constraints across recovered instances within each modality, while the label-based emotion shift unit captures emotional variations in multimodal fusion representations. Experimental results on four benchmark datasets demonstrate that the proposed CRNet achieves remarkable performance compared with state-of-the-art baselines.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| State | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- Emotion recognition in conversation
- incomplete learning
- multimodal learning
Fingerprint
Dive into the research topics of 'CRNet: Context-Guided Recovering Network for Incomplete Multimodal Emotion Recognition in Conversation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver