Skip to main navigation Skip to search Skip to main content

MTF-Net: A mediator transformer-based fusion network with MOE for 6D object pose estimation

  • Zhiqiang Jiang
  • , Tao An
  • , Zimeng Tong
  • , Zhengyu Li
  • , Youran Du
  • , Tao Xie*
  • , Ke Wang
  • , Lijun Zhao
  • , Ruifeng Li
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • Changchun University of Technology
  • Shanghai Ocean University

Research output: Contribution to journalArticlepeer-review

Abstract

6D object pose estimation remains a crucial research field in several computer vision and robotics tasks. Recently, numerous deep learning-based works have demonstrated the feasibility of incorporating RGB and depth images for object pose estimation. However, the primary concerns lie in the optimization of integrating the features of two modalities and addressing dynamic changes of input dominant modality data in various complex scenarios. In this work, we propose MTF-Net, a mediator transformer-based fusion network with adaptive mixture of experts for accurate 6D object pose estimation. For the first concern, we propose a bidirectional mediator transformer (BMT) based fusion module that leverages a linear mediator attention mechanism (LMA) to identify semantic similarities across multi-modality features, thus enabling the network to establish correlations across modalities with reduced computational complexity while maintaining the expressiveness and accuracy of the attention weights. This results in a more effective and potent feature fusion. For the second concern, we introduce an adaptive mixture of experts (A-MOE) layer that can identify the dominant modality data between appearance and geometric features based on the input data, and adjust the network parameters to redistribute weights accordingly, thereby mitigating the impact of low quality data on the pose estimation results. Ultimately, we utilize a 3D keypoint detection network and an instance segmentation module to regress object poses. Comprehensive experiments demonstrate that MTF-Net significantly exceeds state-of-the-art techniques on several benchmarks.

Original languageEnglish
Article number114674
JournalKnowledge-Based Systems
Volume330
DOIs
StatePublished - 25 Nov 2025

Keywords

  • 6D object pose estimation
  • Feature fusion
  • MOE
  • Transformer

Fingerprint

Dive into the research topics of 'MTF-Net: A mediator transformer-based fusion network with MOE for 6D object pose estimation'. Together they form a unique fingerprint.

Cite this