Abstract
Cross-domain object detection in remote sensing suffers from substantial domain gaps arising from differences in resolution, viewing geometry, and imaging modality across sensors and platforms. Existing unsupervised domain-adaptive object detection (DAOD) methods typically align source and target features using the detector’s own target-domain representations. However, the extraction of these representations is constrained by the very domain discrepancies they aim to bridge, resulting in noisy and biased features that make alignment unstable. To address this limitation, we propose the vision–language model-supervised domain adaptor (VLSDA), a domain adaptation framework supervised by a frozen vision–language model (VLM). It leverages a frozen VLM image encoder as an additional and stable semantic domain to guide domain alignment. Our VLM-supervised prototypical alignment (VLPA) module stabilizes categorywise alignment through a tridomain adversarial strategy that jointly aligns source–VLM, target–VLM, and source–target distributions. Complementing this, the global cross-domain contrastive alignment (GCCA) module enhances intraclass compactness and interclass separability via supervised contrastive learning. Without requiring any fine-tuning of the VLM, our framework directly mitigates reliance on noisy target features and improves robustness to large distribution shifts. Extensive experiments on multiple cross-domain remote sensing benchmarks demonstrate consistent improvements over state-of-the-art methods, including 68.3% mAP50 on xView → DOTA and 70.5% on High-Resolution Remote Sensing Detection (HRRSD) dataset → SAR Ship Detection Dataset (SSDD).
| Original language | English |
|---|---|
| Article number | 5651515 |
| Journal | IEEE Transactions on Geoscience and Remote Sensing |
| Volume | 63 |
| DOIs | |
| State | Published - 2025 |
| Externally published | Yes |
Keywords
- Cross-domain alignment
- domain-adaptive object detection (DAOD)
- remote sensing
- vision–language model (VLM)
Fingerprint
Dive into the research topics of 'VLSDA: Vision–Language Model-Supervised Domain Adaptation for Cross-Domain Object Detection in Remote Sensing'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver