Abstract
Dynamic visual odometry is a critical technology for engineering applications including safe autonomous driving and robot navigation. Although traditional geometry-based dynamic visual odometry methods can achieve high accuracy, they often lack robustness in low-texture dynamic environments. In contrast, learning-based methods exhibit better robustness under such conditions but still fall short in accuracy. To bridge this gap, we propose a novel dynamic stereo visual odometry framework that integrates end-to-end learning with multi-frame geometric optimization. Specifically, we design a lightweight Transformer-based network that leverages optical flow and disparity to predict metrically scaled ego-motion. A confidence-guided dynamic mask, derived from both network predictions and semantic priors, enables the network to focus on reliable static regions and suppress the influence of moving objects. Pose estimation is then performed based on a masked attention mechanism. To further enhance accuracy, we introduce a sliding window optimization that refines the initial coarse pose. By leveraging the dynamic masks, we establish reliable matches and optimize the poses using a combination of reprojection factor, an odometry factor, and a scale factor. Extensive experiments on diverse public datasets and a self-collected dataset show that our method achieves state-of-the-art accuracy and robustness in dynamic environments.
| Original language | English |
|---|---|
| Article number | 114611 |
| Journal | Engineering Applications of Artificial Intelligence |
| Volume | 175 |
| DOIs | |
| State | Published - 1 Jul 2026 |
Keywords
- Dynamic environment
- Multi-frame optimization
- Optical flow
- Stereo visual odometry
- Transformer
Fingerprint
Dive into the research topics of 'Towards robust visual odometry in dynamic environments: A hybrid approach with confidence-guided masking'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver