Skip to main navigation Skip to search Skip to main content

VLPRSDet: A vision–language pretrained model for remote sensing object detection

  • School of Electronics and Information Engineering, Harbin Institute of Technology
  • Shanghai Aerospace Electronic Technology Institute

Research output: Contribution to journalArticlepeer-review

Abstract

Recently, numerous excellent vision-language models have emerged in the field of computer vision. These models have demonstrated strong zero-shot detection capabilities and better accuracy after fine-tuning on new datasets in the field of object detection. However, when these models are directly applied to the field of remote sensing, their performance is less than satisfactory. To address this problem, a novel vision-language pretrained model specifically tailored for remote sensing object detection task is proposed. Firstly, we create a new dataset composed of object-text pairs by collecting a large amount of remote sensing image object detection data to train the proposed model. Then, by integrating the CLIP model in the field of remote sensing with the YOLO detector, we propose a vision-language pretrained model for remote sensing object detection (VLPRSDet). VLPRSDet achieves enhanced fusion of visual and textual features through a vision language path aggregation network, and then aligns visual embeddings and textual embeddings through Region Text Matching to achieve the alignment between object regions and text. Experimental results indicate that the proposed VLPRSDet exhibits robust zero-shot capabilities in the field of remote sensing object detection, and can achieve superior detection accuracy after fine-tuning on specific datasets. Specifically, after fine-tuning, VLPRSDet can achieve 76.2 % mAP on the DIOR dataset and 94.2 % mAP on the HRRSD dataset. The code and dataset will be released at https://github.com/dyl96/VLPRSDet.

Original languageEnglish
Article number131712
JournalNeurocomputing
Volume658
DOIs
StatePublished - 28 Dec 2025

Keywords

  • Pretrained
  • Remote sensing object detection
  • Vision-language model
  • Zero-shot

Fingerprint

Dive into the research topics of 'VLPRSDet: A vision–language pretrained model for remote sensing object detection'. Together they form a unique fingerprint.

Cite this