Skip to main navigation Skip to search Skip to main content

M2F2: A Multi-scale and Multi-view Feature Fusion Network for Generalizable Object Pose Estimation

  • Hongbo Gao
  • , Zhengyu Li
  • , Mengyuan Wu
  • , Tao Xie*
  • , Ruifeng Li*
  • , Zhendong Fan
  • , Kun Dai
  • , Xin Wen
  • , Lijun Zhao
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • Yangtze River Delta HIT Robot Technology Research Institute
  • Harbin institute of technology

Research output: Contribution to journalArticlepeer-review

Abstract

Object pose estimation is a fundamental task in computer vision and robotics, with recent progress driven by generalizable model-free paradigms such as Gen6D, which typically consist of an object detector, a viewpoint selector, and a pose refiner. Despite their promising performance on unseen objects, existing methods still suffer from limited generalization due to suboptimal feature representation and fusion mechanisms. Specifically, the use of pre-trained visual encoders in the object detector often fails to capture task-specific structural and semantic cues. More critically, current multi-view matching strategies implicitly rely on globally consistent aggregation, overlooking the spatially varying reliability of different reference views, which leads to mismatched correspondences and degraded localization accuracy. In addition, naive multi-scale feature fusion in the view-point selector neglects the varying importance of features across spatial and channel dimensions. To address these limitations, we propose M2F2, a novel Gen6D-based framework with enhanced feature compensation and fusion. First, a co-information guided feature compensation module (CFCM) is introduced to improve the extraction of shared semantic and geometric information between query and reference images. Second, we design an element-wise viewpoint fusion module (VFM) that explicitly models spatially varying reliability across reference views and performs selective fusion via element-wise top-K selection and adaptive weighting, enabling spatially adaptive correspondences. Third, a scale fusion module (SFM) is proposed to effectively integrate multi-scale features by jointly modeling spatial and channel-wise dependencies. Extensive experiments on GenMOP and LINEMOD demonstrate that M2F2 achieves superior performance over state-of-the-art methods, highlighting its effectiveness and generalization capability.

Keywords

  • Feature compensation
  • Multi-scale feature fusion
  • Multi-view feature fusion
  • Object pose estimation

Fingerprint

Dive into the research topics of 'M2F2: A Multi-scale and Multi-view Feature Fusion Network for Generalizable Object Pose Estimation'. Together they form a unique fingerprint.

Cite this