Skip to main navigation Skip to search Skip to main content

Echo: Generating cross-modal features for unseen classes in zero-shot remote sensing image captioning

  • Kangda Cheng
  • , Jinlong Liu*
  • , Rui Mao
  • , Zhilu Wu
  • , Erik Cambria
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • Nanyang Technological University

Research output: Contribution to journalArticlepeer-review

Abstract

Remote Sensing Image Captioning (RSIC) has broad application prospects in multimodal information fusion and remote sensing data analysis. However, traditional methods heavily rely on large-scale labeled data, making it challenging to generate captions in zero-shot scenarios. To the best of our knowledge, there is no research specifically addressing zero-shot RSIC (ZS-RSIC). Zero-shot image captioning methods for natural images typically achieve their goals by constructing a shared intermediate semantic space that connects the visual space and the category space. However, these approaches often exhibit a strong bias towards seen classes. Furthermore, due to the systematic discrepancy between natural image priors and remote sensing specific semantics, coupled with the lack of an effective mapping mechanism between pixel-level details and professional terminology captions, the fineness of captions generated by large-scale pre-training methods and the understanding of scene context tend to be insufficient. To address these challenges, we propose Echo, which generates synthetic cross-modal features for these unseen classes in the ZS-RSIC task. The synthetic features closely approximate the distribution and semantics of real unseen classes by leveraging cross-modal feature generation and semantic alignment techniques. This approach reframes the zero-shot learning task as an approximated supervised learning problem within the feature space. We also propose a pseudo-caption self-training strategy, which iteratively self-trains by selecting high-confidence pseudo-samples, thereby gradually correcting the bias of the feature generator. This strategy is designed to further mitigate the domain shift problem.

Original languageEnglish
Article number103952
JournalInformation Fusion
Volume128
DOIs
StatePublished - Apr 2026

Keywords

  • Feature generation
  • Image captioning
  • Remote sensing
  • Zero-shot learning

Fingerprint

Dive into the research topics of 'Echo: Generating cross-modal features for unseen classes in zero-shot remote sensing image captioning'. Together they form a unique fingerprint.

Cite this