Skip to main navigation Skip to search Skip to main content

Real-time Stereo Matching Networks based on Multi-modal and Sparse-dense Fusion

  • Xi Zhang
  • , Xiaojun Wu*
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Binocular depth estimation models using deep learning struggle to balance accuracy and efficiency, consuming significant memory during inference. We present a novel, concise, and scalable model architecture, Fusion-Stereo, based on disparity iterative optimization. It ensures excellent prediction accuracy for stereo vision sensors while achieving real-time inference speed and low memory usage. Observing the different modal information of image texture, matching similarity, and disparity features, we propose a new disparity residual prediction module for adaptive propagation of reliable disparity information in spatial and disparity dimensions. To enhance the reliability of depth prediction, an efficient fusion mechanism is designed between Fusion-Stereo and external prior data sources. Sparse disparity priors are introduced to the model to assist dense disparity prediction, maintaining essentially unchanged inference speed and memory occupation.

Original languageEnglish
Title of host publicationRCAR 2025 - IEEE International Conference on Real-Time Computing and Robotics
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages1198-1203
Number of pages6
ISBN (Electronic)9798331502058
DOIs
StatePublished - 2025
Externally publishedYes
Event2025 IEEE International Conference on Real-Time Computing and Robotics, RCAR 2025 - Toyama, Japan
Duration: 1 Jun 20256 Jun 2025

Publication series

NameRCAR 2025 - IEEE International Conference on Real-Time Computing and Robotics

Conference

Conference2025 IEEE International Conference on Real-Time Computing and Robotics, RCAR 2025
Country/TerritoryJapan
CityToyama
Period1/06/256/06/25

Fingerprint

Dive into the research topics of 'Real-time Stereo Matching Networks based on Multi-modal and Sparse-dense Fusion'. Together they form a unique fingerprint.

Cite this