Skip to main navigation Skip to search Skip to main content

Toward Multimodal Sentiment Analysis via Contrastive Cross-Modal Retrieval Augmentation and Hierachical Prompts

  • Xianbing Zhao
  • , Shengzun Yang
  • , Buzhou Tang*
  • , Ronghuan Jiang*
  • *Corresponding author for this work
  • Jiangnan University
  • Harbin Institute of Technology
  • Guangdong Key Laboratory of Intelligent Transportation Systems
  • Pengcheng Laboratory
  • General Hospital of People's Liberation Army

Research output: Contribution to journalArticlepeer-review

Abstract

Multimodal Sentiment Analysis (MSA) is a fundamental problem in the field of affective computing. Although significant progress has been made in cross-modal interaction, it remains a challenge due to the insufficient reference context in cross-modal interactions. Current cross-modal approaches primarily focus on leveraging modality-level reference context within a individual sample for cross-modal feature enhancement, neglecting the potential cross-sample relationships that can serve as sample-level reference context to enhance the cross-modal features. To address this issue, we propose a novel multimodal retrieval-augmented framework to simultaneously incorporate cross-sample modality-level reference context and cross-sample sample-level reference context to enhance the multimodal features. In particular, we first design a contrastive cross-modal retrieval module to retrieve semantic similar samples and enhance anchor modality. To endow the model to capture both cross-sample and intra-sample information, we integrate two different types of prompts, modality-level prompts and sample-level prompts, to generate modality-level and sample-level reference contexts, respectively. Finally, we design a cross-modal retrieval-augmented encoder that simultaneously leverages modality-level and sample-level reference contexts to enhance the anchor modality. Extensive experiments demonstrate the effectiveness and superiority of our model on two publicly available datasets.

Original languageEnglish
Pages (from-to)2091-2104
Number of pages14
JournalIEEE Transactions on Affective Computing
Volume17
Issue number2
DOIs
StatePublished - Apr 2026
Externally publishedYes

Keywords

  • Multimodal sentiment analysis
  • multimodal retrieval augmentation
  • prompt learning

Fingerprint

Dive into the research topics of 'Toward Multimodal Sentiment Analysis via Contrastive Cross-Modal Retrieval Augmentation and Hierachical Prompts'. Together they form a unique fingerprint.

Cite this