Abstract
Three-dimensional object detection from point clouds represents a formidable challenge, necessitating the accurate identification and localization of objects within a 3-D space. Recent advancements have showcased the efficacy of point-based detectors, leveraging local aggregators to encode intricate structural details of the point cloud. However, a notable limitation resides in their treatment of each point and object proposal in isolation, devoid of considering the interrelationships among them, thus impeding the overall detection performance. In this work, we argue that the integration of contextual information is paramount, particularly in the realm of indoor 3-D object detection. Indoor environments are inherently characterized by robust contextual constraints, providing a rich tapestry for enhanced scene comprehension. In this way, we introduce MTNet, a mixed transformer (MixFormer) network for high-quality indoor 3-D object detection. Technically, we develop a MixFormer block that is purpose-built to intricately model the synergistic interplay between local structural information and global contextual features of 3-D point clouds. In contrast to the classical transformer, our proposed MixFormer incorporates a local feature aggregator (LFA) engineered to capture local geometric information while leveraging a K-nearest neighbors (KNNs)-based attention mechanism to aggregate global contextual information. Furthermore, we suggest a feature compensator to adaptively fuse the strengths of local and global features, further bolstering detection performance. The culmination of these proposed components results in our MTNet framework, a hierarchical, versatile pipeline that consistently outperforms existing works across a multitude of benchmarks. In addition, we affirm the potential of the proposed MixFormer and feature compensator (FC) as generic modules that are capable of augmenting performance across a spectrum of 3-D downstream point cloud tasks.
| Original language | English |
|---|---|
| Pages (from-to) | 28197-28211 |
| Number of pages | 15 |
| Journal | IEEE Internet of Things Journal |
| Volume | 13 |
| Issue number | 13 |
| DOIs | |
| State | Published - 1 Jul 2026 |
Keywords
- 3-D object detection
- deep learning
- transformer
Fingerprint
Dive into the research topics of 'MTNet: A Mixed Transformer Network for High-Quality 3-D Object Detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver