Abstract
Current vision-based spatial pose estimation methods primarily rely on CNNs for feature extraction. However, CNN-based approaches struggle to model the global relationships within an image, leading to suboptimal performance when handling variations in target scale and complex spatial backgrounds. This paper proposes a novel spatial pose estimation method for non-cooperative targets based on a vision transformer. Specifically, we design a keypoint position regression network that utilizes Swin Transformer to extract multi-level features from target images, capturing both global structural information and fine texture details. To adapt to variations in target scale, we fuse feature maps at different resolutions to construct a feature pyramid. Finally, a series of convolutional modules regress the target keypoint positions. By combining the estimated 3D coordinates of keypoints with the EPNP algorithm, we obtain the final target pose. Experimental results demonstrate that the proposed method achieves high-precision pose estimation even under varying target scales and complex background conditions.
| Original language | English |
|---|---|
| Pages (from-to) | 309-314 |
| Number of pages | 6 |
| Journal | IFAC-PapersOnLine |
| Volume | 59 |
| Issue number | 20 |
| DOIs | |
| State | Published - 1 Aug 2025 |
| Externally published | Yes |
| Event | 23th IFAC Symposium on Automatic Control in Aerospace, ACA 2025 - Harbin, China Duration: 2 Aug 2025 → 6 Aug 2025 |
Keywords
- Keypoint detection
- monocular vision
- non-cooperative targets
- pose estimation
- vision Transformer
Fingerprint
Dive into the research topics of 'SwinFPN: A Pose Estimation Method Adapted to Multi-Scale Variations of Space Targets'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver