Skip to main navigation Skip to search Skip to main content

ViSP: A PPO-enhanced framework for multimodal sarcasm generation with contrastive learning

  • Changli Wang
  • , Fang Yin
  • , Jiafeng Liu
  • , Rui Wu*
  • *Corresponding author for this work
  • Faculty of Computing, Harbin Institute of Technology
  • Harbin University of Science and Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Human emotions are inherently complex, with sarcasm representing one of the most subtle and distinctive forms. Although significant progress has been made in sarcasm understanding, sarcasm generation remains underexplored, largely due to the overreliance on textual modalities, the neglect of visual cues, and the semantic mismatch between images and sarcastic intent in existing datasets. In this work, we present M2SaG, a multimodal sarcasm generation dataset comprising 4970 samples, each containing an image, a sarcastic text, and its corresponding sarcasm target. To benchmark M2SaG, we propose ViSP, a ViLT-based sarcasm generation framework that integrates Proximal Policy Optimization (PPO) with contrastive learning. PPO leverages reward scores derived from DIP to guide the generation process, while contrastive learning encourages the model to prefer outputs with higher rewards. These strategies jointly enhance generation quality and promote stronger sarcastic intent in the outputs. Comprehensive evaluations across five metric sets show that ViSP consistently outperforms all baseline models, including Large Language Models, highlighting their limitations in sarcasm generation. Furthermore, analysis of Sarcasm Scores and Factual Incongruity distributions for both M2SaG and ViSP outputs reveals that ViSP achieves higher mean Sarcasm Scores (0.898 vs. 0.770) and Factual Incongruity (0.768 vs. 0.739), demonstrating its ability to produce higher-quality and more contextually sarcastic texts. Our dataset is available at https://github.com/wclapply/ViSP.

Original languageEnglish
Article number133185
JournalNeurocomputing
Volume678
DOIs
StatePublished - 14 May 2026
Externally publishedYes

Keywords

  • Contrastive learning
  • Proximal policy optimization
  • Sarcasm generation
  • Vision-language generation

Fingerprint

Dive into the research topics of 'ViSP: A PPO-enhanced framework for multimodal sarcasm generation with contrastive learning'. Together they form a unique fingerprint.

Cite this