Skip to main navigation Skip to search Skip to main content

DMFF: dual-way multimodal feature fusion for 3D object detection

  • Xiaopeng Dong
  • , Xiaoguang Di*
  • , Wenzhuang Wang
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Recently, multimodal 3D object detection that fuses the complementary information from LiDAR data and RGB images has been an active research topic. However, it is not trivial to fuse images and point clouds because of different representations of them. Inadequate feature fusion also brings bad effects on detection performance. We convert images into pseudo point clouds by using a depth completion and utilize a more efficient feature fusion method to address the problems. In this paper, we propose a dual-way multimodal feature fusion network (DMFF) for 3D object detection. Specifically, we first use a dual stream feature extraction module (DSFE) to generate homogeneous LiDAR and pseudo region of interest (RoI) features. Then, we propose a dual-way feature interaction method (DWFI) that enables intermodal and intramodal interaction of the two features. Next, we design a local attention feature fusion module (LAFF) to select which features of the input are more likely to contribute to the desired output. In addition, the proposed DMFF achieves the state-of-the-art performances on the KITTI Dataset.

Original languageEnglish
Pages (from-to)455-463
Number of pages9
JournalSignal, Image and Video Processing
Volume18
Issue number1
DOIs
StatePublished - Feb 2024

Keywords

  • 3D object detection
  • Lidar point clouds
  • Multimodal feature fusion
  • RGB images
  • Self-attention mechanism

Fingerprint

Dive into the research topics of 'DMFF: dual-way multimodal feature fusion for 3D object detection'. Together they form a unique fingerprint.

Cite this