Skip to main navigation Skip to search Skip to main content

SwinFPN: A Pose Estimation Method Adapted to Multi-Scale Variations of Space Targets

  • Fan Zhang*
  • , Zexu Zhang*
  • , Zhuo Song*
  • , Yicheng Mao*
  • *Corresponding author for this work
  • School of Astronautics, Harbin Institute of Technology

Research output: Contribution to journalConference articlepeer-review

Abstract

Current vision-based spatial pose estimation methods primarily rely on CNNs for feature extraction. However, CNN-based approaches struggle to model the global relationships within an image, leading to suboptimal performance when handling variations in target scale and complex spatial backgrounds. This paper proposes a novel spatial pose estimation method for non-cooperative targets based on a vision transformer. Specifically, we design a keypoint position regression network that utilizes Swin Transformer to extract multi-level features from target images, capturing both global structural information and fine texture details. To adapt to variations in target scale, we fuse feature maps at different resolutions to construct a feature pyramid. Finally, a series of convolutional modules regress the target keypoint positions. By combining the estimated 3D coordinates of keypoints with the EPNP algorithm, we obtain the final target pose. Experimental results demonstrate that the proposed method achieves high-precision pose estimation even under varying target scales and complex background conditions.

Original languageEnglish
Pages (from-to)309-314
Number of pages6
JournalIFAC-PapersOnLine
Volume59
Issue number20
DOIs
StatePublished - 1 Aug 2025
Externally publishedYes
Event23th IFAC Symposium on Automatic Control in Aerospace, ACA 2025 - Harbin, China
Duration: 2 Aug 20256 Aug 2025

Keywords

  • Keypoint detection
  • monocular vision
  • non-cooperative targets
  • pose estimation
  • vision Transformer

Fingerprint

Dive into the research topics of 'SwinFPN: A Pose Estimation Method Adapted to Multi-Scale Variations of Space Targets'. Together they form a unique fingerprint.

Cite this