Abstract
Highlights: What are the main findings? LiteSAM, a lightweight satellite–UAV feature matching framework, employs unified feature representation and fine-grained matching strategies to achieve a superior trade-off between accuracy and efficiency in cross-view matching. Achieves state-of-the-art performance on multiple benchmarks while substantially reducing model complexity and inference latency. What is the implication of the main finding? Enables real-time UAV visual localization in GPS-denied and resource-constrained environments, making it practical for deployment. Demonstrates strong generalization across datasets and scenarios, extending its applicability to remote sensing and natural image matching tasks. We present a (Light)weight (S)atellite–(A)erial feature (M)atching framework (LiteSAM) for robust UAV absolute visual localization (AVL) in GPS-denied environments. Existing satellite–aerial matching methods struggle with large appearance variations, texture-scarce regions, and limited efficiency for real-time UAV applications. LiteSAM integrates three key components to address these issues. First, efficient multi-scale feature extraction optimizes representation, reducing inference latency for edge devices. Second, a Token Aggregation–Interaction Transformer (TAIFormer) with a convolutional token mixer (CTM) models inter- and intra-image correlations, enabling robust global–local feature fusion. Third, a MinGRU-based dynamic subpixel refinement module adaptively learns spatial offsets, enhancing subpixel-level matching accuracy and cross-scenario generalization. The experiments show that LiteSAM achieves competitive performance across multiple datasets. On UAV-VisLoc, LiteSAM attains an RMSE@30 of 17.86 m, outperforming state-of-the-art semi-dense methods such as EfficientLoFTR. Its optimized variant, LiteSAM (opt., without dual softmax), delivers inference times of 61.98 ms on standard GPUs and 497.49 ms on NVIDIA Jetson AGX Orin, which are 22.9% and 19.8% faster than EfficientLoFTR (opt.), respectively. With 6.31M parameters, which is 2.4× fewer than EfficientLoFTR’s 15.05M, LiteSAM proves to be suitable for edge deployment. Extensive evaluations on natural image matching and downstream vision tasks confirm its superior accuracy and efficiency for general feature matching.
| Original language | English |
|---|---|
| Article number | 3349 |
| Journal | Remote Sensing |
| Volume | 17 |
| Issue number | 19 |
| DOIs | |
| State | Published - Oct 2025 |
| Externally published | Yes |
Keywords
- absolute visual localization
- convolutional token mixer
- feature matching
- satellite–aerial imagery matching
Fingerprint
Dive into the research topics of 'LiteSAM: Lightweight and Robust Feature Matching for Satellite and Aerial Imagery'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver