Skip to main navigation Skip to search Skip to main content

CMCap: Cross-Modal Meta-Captioning for Few-Shot Remote Sensing Image Captioning

  • Kangda Cheng
  • , Jinlong Liu*
  • , Rui Mao
  • , Zhilu Wu
  • , Erik Cambria
  • *Corresponding author for this work
  • School of Electronics and Information Engineering, Harbin Institute of Technology
  • Nanyang Technological University

Research output: Contribution to journalArticlepeer-review

Abstract

Remote sensing image captioning (RSIC) aims to automatically generate natural language descriptions for remote sensing images, which holds significant application value. However, existing methods heavily rely on large-scale annotated datasets. In practical applications, due to the high annotation cost of remote sensing data and the dynamic diversity of ground object categories, only a limited number of labeled samples are available for many scenarios, leading to severe data scarcity challenges. To address this challenge, we propose the cross-modal meta-captioning (CMCap) framework. First, we design a sufficient task sampling (STS) strategy that constructs a large number of pseudo-tasks in each iteration to reduce gradient variance and improve data utilization. Then, we synthesize challenging support samples to enhance discriminative representation learning, and propose a task-adaptive cross-modal alignment (TCA) module to decouple task adaptability from general representation learning, thereby improving the model's discriminability and cross-task generalization under extremely few-shot conditions. Moreover, we develop a meta-conditioning decoder that deeply fuses task context with query image features to guide a frozen large language model (LLM) in generating task-consistent captions. Experiments conducted on UCM-Captions, Sydney-Captions, and RSICD demonstrate that CMCap achieves competitive performance across various evaluation metrics and exhibits remarkable generalization capability under extreme data scarcity conditions.

Original languageEnglish
Article number6006705
JournalIEEE Geoscience and Remote Sensing Letters
Volume23
DOIs
StatePublished - 2026
Externally publishedYes

Keywords

  • Few-shot learning
  • image captioning
  • meta-learning
  • remote sensing

Fingerprint

Dive into the research topics of 'CMCap: Cross-Modal Meta-Captioning for Few-Shot Remote Sensing Image Captioning'. Together they form a unique fingerprint.

Cite this