Skip to main navigation Skip to search Skip to main content

DTF-Net: A Dual Attention Transformer-based Fusion Network for 6D Object Pose Estimation

  • Tao An
  • , Kun Dai
  • , Ruifeng Li*
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In the field of computer vision, 6D object pose estimation has emerged as a crucial research direction. Recent advancements leverage deep learning networks that integrate RGB images with depth images for enhanced feature learning. In this study, we introduce DTF-Net, a dual attention transformer-based fusion network for 6D object pose estimation using RGBD data. By combining appearance features from RGB data with geometric features from depth images, our method comprehensively captures object characteristics in a scene, significantly enhancing pose estimation accuracy. Central to our approach is the dual attention transformer (DAT) based fusion module, which fuses these two complementary features. The DAT surpasses traditional transformer by accurately capturing dual-modal data features and reducing computational complexity. Our network's superior performance is validated through experiments on the YCB-Video dataset, where it outperforms current state-of-the-art models.

Original languageEnglish
Title of host publicationProceedings - 2024 16th International Conference on Intelligent Human-Machine Systems and Cybernetics, IHMSC 2024
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages104-107
Number of pages4
ISBN (Electronic)9798331540241
DOIs
StatePublished - 2024
Event16th International Conference on Intelligent Human-Machine Systems and Cybernetics, IHMSC 2024 - Hangzhou, China
Duration: 24 Aug 202425 Aug 2024

Publication series

NameProceedings - 2024 16th International Conference on Intelligent Human-Machine Systems and Cybernetics, IHMSC 2024

Conference

Conference16th International Conference on Intelligent Human-Machine Systems and Cybernetics, IHMSC 2024
Country/TerritoryChina
CityHangzhou
Period24/08/2425/08/24

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • Deep learning
  • Fusion feature
  • Pose estimation
  • RGBD

Fingerprint

Dive into the research topics of 'DTF-Net: A Dual Attention Transformer-based Fusion Network for 6D Object Pose Estimation'. Together they form a unique fingerprint.

Cite this