Abstract
6D object pose estimation remains a crucial research field in several computer vision and robotics tasks. Recently, numerous deep learning-based works have demonstrated the feasibility of incorporating RGB and depth images for object pose estimation. However, the primary concerns lie in the optimization of integrating the features of two modalities and addressing dynamic changes of input dominant modality data in various complex scenarios. In this work, we propose MTF-Net, a mediator transformer-based fusion network with adaptive mixture of experts for accurate 6D object pose estimation. For the first concern, we propose a bidirectional mediator transformer (BMT) based fusion module that leverages a linear mediator attention mechanism (LMA) to identify semantic similarities across multi-modality features, thus enabling the network to establish correlations across modalities with reduced computational complexity while maintaining the expressiveness and accuracy of the attention weights. This results in a more effective and potent feature fusion. For the second concern, we introduce an adaptive mixture of experts (A-MOE) layer that can identify the dominant modality data between appearance and geometric features based on the input data, and adjust the network parameters to redistribute weights accordingly, thereby mitigating the impact of low quality data on the pose estimation results. Ultimately, we utilize a 3D keypoint detection network and an instance segmentation module to regress object poses. Comprehensive experiments demonstrate that MTF-Net significantly exceeds state-of-the-art techniques on several benchmarks.
| Original language | English |
|---|---|
| Article number | 114674 |
| Journal | Knowledge-Based Systems |
| Volume | 330 |
| DOIs | |
| State | Published - 25 Nov 2025 |
Keywords
- 6D object pose estimation
- Feature fusion
- MOE
- Transformer
Fingerprint
Dive into the research topics of 'MTF-Net: A mediator transformer-based fusion network with MOE for 6D object pose estimation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver