Skip to main navigation Skip to search Skip to main content

CADET: Cycle-Consistent Adversarial Disentanglement With Text-Guided Enhancement for Multimodal Sentiment Analysis

  • Xiangrui Dan
  • , Bin Gao*
  • , Yanping Chen
  • , Qian Fu
  • , Yutong Li
  • , Shutian Liu
  • , Zhengjun Liu
  • *Corresponding author for this work
  • Heilongjiang University
  • Guizhou University
  • School of Physics, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Multimodal sentiment analysis (MSA) aims to predict sentiment from language, vision, and audio modalities. Although these modalities provide complementary affective cues, they contribute unequally to sentiment understanding and exhibit distinct expressive patterns. Existing methods often strengthen cross-modal interaction or directly fuse heterogeneous features, but they usually pay limited attention to the structure of pre-fusion representations. Consequently, shared sentiment information may become entangled with modality-specific cues, making these complementary cues difficult to preserve during multimodal fusion. To address these challenges, we propose CADET, a cycle-consistent adversarial disentanglement framework with text-guided enhancement. The proposed framework decomposes each modality into shared and specific representations, organizes the resulting latent spaces through structured regularization, and employs cycle-consistent reconstruction to preserve semantic consistency during cross-modal transformation. Since language usually provides the most explicit sentiment information, CADET uses textual representations to refine sentiment-relevant modality-specific features in the visual and audio branches through cross-modal attention. Extensive experiments on three public benchmarks, CMU-MOSI, CMU-MOSEI, and CH-SIMS, demonstrate that CADET achieves superior performance over state-of-the-art methods. These results show the effectiveness of combining disentanglement, semantic preservation, and text-guided enhancement for multimodal sentiment prediction.

Original languageEnglish
JournalIEEE Transactions on Affective Computing
DOIs
StateAccepted/In press - 2026
Externally publishedYes

Keywords

  • Cross-modal representation learning
  • cycle consistency
  • multimodal sentiment analysis
  • text-guided enhancement

Fingerprint

Dive into the research topics of 'CADET: Cycle-Consistent Adversarial Disentanglement With Text-Guided Enhancement for Multimodal Sentiment Analysis'. Together they form a unique fingerprint.

Cite this