Abstract
Multimodal sentiment analysis (MSA) aims to predict sentiment from language, vision, and audio modalities. Although these modalities provide complementary affective cues, they contribute unequally to sentiment understanding and exhibit distinct expressive patterns. Existing methods often strengthen cross-modal interaction or directly fuse heterogeneous features, but they usually pay limited attention to the structure of pre-fusion representations. Consequently, shared sentiment information may become entangled with modality-specific cues, making these complementary cues difficult to preserve during multimodal fusion. To address these challenges, we propose CADET, a cycle-consistent adversarial disentanglement framework with text-guided enhancement. The proposed framework decomposes each modality into shared and specific representations, organizes the resulting latent spaces through structured regularization, and employs cycle-consistent reconstruction to preserve semantic consistency during cross-modal transformation. Since language usually provides the most explicit sentiment information, CADET uses textual representations to refine sentiment-relevant modality-specific features in the visual and audio branches through cross-modal attention. Extensive experiments on three public benchmarks, CMU-MOSI, CMU-MOSEI, and CH-SIMS, demonstrate that CADET achieves superior performance over state-of-the-art methods. These results show the effectiveness of combining disentanglement, semantic preservation, and text-guided enhancement for multimodal sentiment prediction.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Affective Computing |
| DOIs | |
| State | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- Cross-modal representation learning
- cycle consistency
- multimodal sentiment analysis
- text-guided enhancement
Fingerprint
Dive into the research topics of 'CADET: Cycle-Consistent Adversarial Disentanglement With Text-Guided Enhancement for Multimodal Sentiment Analysis'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver