Skip to main navigation Skip to search Skip to main content

Co-Training Vision-Language Models for Remote Sensing Multi-Task Learning †

  • Qingyun Li
  • , Shuran Ma
  • , Junwei Luo
  • , Yi Yu
  • , Yue Zhou
  • , Fengxiang Wang
  • , Xudong Lu
  • , Xiaoxing Wang
  • , Xin He
  • , Yushi Chen*
  • , Xue Yang
  • *Corresponding author for this work
  • School of Electronics and Information Engineering, Harbin Institute of Technology
  • School of Telecommunications Engineering, Xidian University
  • Wuhan University
  • Southeast University, Nanjing
  • East China Normal University
  • National University of Defense Technology
  • Chinese University of Hong Kong
  • Shanghai Jiao Tong University

Research output: Contribution to journalArticlepeer-review

Abstract

Highlights: What are the main findings? The proposed RSCoVLM, trained on a well-curated data recipe, is a fully open-sourced vision-language model (VLM) that excels in multiple remote sensing (RS) tasks. It achieves state-of-the-art performance across various tasks, even aerial object detection. The proposed dynamic resolution strategies enable the processing of RS images of arbitrary sizes. Among these strategies, the Zoom-in Chain method significantly enhances the performance on ultra-high-resolution RS images reasoning. Additionally, VLMs exhibit clear limitations under the commonly used mAP metric, which is influenced by confidence scores. Based on the proposed (Formula presented.), RSCoVLM is shown to achieve detection performance comparable to conventional object detection models. What are the implication of the main findings? From the perspective of RS VLM development, as a new baseline, RSCoVLM demonstrates substantial progress in capability and flexibility. It brings us one step closer to realizing a general-purpose generative agent for RS image processing. From the perspective of RS multi-task learning, the proposed framework offers greater extensibility. It will facilitate expansion to more and increasingly complex tasks in the future, moving toward a unified multi-task model. With Transformers achieving outstanding performance on individual remote sensing (RS) tasks, we are now approaching the realization of a unified model that excels across multiple tasks through multi-task learning (MTL). Compared to single-task approaches, MTL methods offer improved generalization, enhanced scalability, and greater practical applicability. Recently, vision-language models (VLMs) have achieved promising results in RS image understanding, grounding, and ultra-high-resolution (UHR) image reasoning, respectively. Moreover, the unified text-based interface demonstrates significant potential for MTL. Hence, in this work, we present RSCoVLM, a simple yet flexible VLM baseline for RS MTL. Firstly, we create the data curation procedure, including data acquisition, offline processing and integrating, as well as online loading and weighting. This data procedure effectively addresses complex RS data enviroments and generates flexible vision-language conversations. Furthermore, we propose a unified dynamic-resolution strategy to address the diverse image scales inherent in RS imagery. For UHR images, we introduce the Zoom-in Chain mechanism together with its corresponding dataset, LRS-VQA-Zoom. The strategies are flexible and effectively mitigate the computational burdens. Additionally, we significantly enhance the model’s object detection capability and propose a novel evaluation protocol that ensures fair comparison between VLMs and conventional detection models. Extensive experiments demonstrate that RSCoVLM achieves state-of-the-art performance across diverse tasks, outperforming existing RS VLMs and even rivaling specialized expert models. All the training and evaluating tools, model weights, and datasets have been fully open-sourced to support reproducibility. We expect that this baseline will promote further progress toward general-purpose RS models.

Original languageEnglish
Article number222
JournalRemote Sensing
Volume18
Issue number2
DOIs
StatePublished - Jan 2026
Externally publishedYes

Keywords

  • multi-task learning
  • remote sensing
  • vision-language model

Fingerprint

Dive into the research topics of 'Co-Training Vision-Language Models for Remote Sensing Multi-Task Learning †'. Together they form a unique fingerprint.

Cite this