Skip to main navigation Skip to search Skip to main content

Toward Comprehensive Interactive Change Understanding in Remote Sensing: A Large-Scale Dataset and Dual-Granularity Enhanced VLM

  • Junxiao Xue
  • , Quan Deng
  • , Xuecheng Wu
  • , Kelu Yao
  • , Xinyi Yin
  • , Fei Yu
  • , Wei Zhou
  • , Yanfei Zhong
  • , Yang Liu*
  • , Dingkang Yang*
  • *Corresponding author for this work
  • Zhejiang Lab
  • University of Chinese Academy of Sciences
  • Xi'an Jiaotong University
  • Zhengzhou University
  • Liaoning University of Technology
  • Cardiff University
  • Wuhan University
  • Tongji University
  • Fudan University
  • Fysics AI

Research output: Contribution to journalArticlepeer-review

Abstract

Remote sensing change understanding (RSCU) is essential for analyzing remote sensing images and understanding how human activities affect the environment. However, existing datasets lack a deep understanding and interactions in the diverse change captioning, counting, and localization tasks. To tackle these gaps, we construct ChangeIMTI, a new large-scale interactive multitask instruction dataset that encompasses four complementary tasks, including change captioning, binary change classification, change counting, and change localization. Building upon this new dataset, we further design a novel vision-guided vision–language model (ChangeVG) with dual-granularity awareness for bitemporal remote sensing images (i.e., two remote sensing images of the same area at different times). The introduced vision-guided module is a dual-branch architecture that synergistically combines fine-grained spatial feature extraction with high-level semantic summarization. These enriched representations further serve as the auxiliary prompts to guide large vision–language models (VLMs) (e.g., Qwen2.5-VL-7B) during instruction tuning, thereby facilitating the hierarchical cross-modal learning. We extensively conduct experiments across four tasks to demonstrate the superiority of our approach. Remarkably, on the change captioning task, our method outperforms the strongest method Semantic-CC by 1.39 points on the comprehensive S*m metric, which integrates the semantic similarity and descriptive accuracy to provide an overall evaluation of change caption. Moreover, we also perform a series of ablation studies to examine the critical components of our method.

Original languageEnglish
Article number4401516
JournalIEEE Transactions on Geoscience and Remote Sensing
Volume64
DOIs
StatePublished - 2026
Externally publishedYes

Keywords

  • Change captioning
  • dataset
  • large vision–language models (VLMs)
  • remote sensing change understanding (RSCU)

Fingerprint

Dive into the research topics of 'Toward Comprehensive Interactive Change Understanding in Remote Sensing: A Large-Scale Dataset and Dual-Granularity Enhanced VLM'. Together they form a unique fingerprint.

Cite this