TY - GEN
T1 - Real-time Stereo Matching Networks based on Multi-modal and Sparse-dense Fusion
AU - Zhang, Xi
AU - Wu, Xiaojun
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Binocular depth estimation models using deep learning struggle to balance accuracy and efficiency, consuming significant memory during inference. We present a novel, concise, and scalable model architecture, Fusion-Stereo, based on disparity iterative optimization. It ensures excellent prediction accuracy for stereo vision sensors while achieving real-time inference speed and low memory usage. Observing the different modal information of image texture, matching similarity, and disparity features, we propose a new disparity residual prediction module for adaptive propagation of reliable disparity information in spatial and disparity dimensions. To enhance the reliability of depth prediction, an efficient fusion mechanism is designed between Fusion-Stereo and external prior data sources. Sparse disparity priors are introduced to the model to assist dense disparity prediction, maintaining essentially unchanged inference speed and memory occupation.
AB - Binocular depth estimation models using deep learning struggle to balance accuracy and efficiency, consuming significant memory during inference. We present a novel, concise, and scalable model architecture, Fusion-Stereo, based on disparity iterative optimization. It ensures excellent prediction accuracy for stereo vision sensors while achieving real-time inference speed and low memory usage. Observing the different modal information of image texture, matching similarity, and disparity features, we propose a new disparity residual prediction module for adaptive propagation of reliable disparity information in spatial and disparity dimensions. To enhance the reliability of depth prediction, an efficient fusion mechanism is designed between Fusion-Stereo and external prior data sources. Sparse disparity priors are introduced to the model to assist dense disparity prediction, maintaining essentially unchanged inference speed and memory occupation.
UR - https://www.scopus.com/pages/publications/105016844354
U2 - 10.1109/RCAR65431.2025.11139590
DO - 10.1109/RCAR65431.2025.11139590
M3 - 会议稿件
AN - SCOPUS:105016844354
T3 - RCAR 2025 - IEEE International Conference on Real-Time Computing and Robotics
SP - 1198
EP - 1203
BT - RCAR 2025 - IEEE International Conference on Real-Time Computing and Robotics
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 IEEE International Conference on Real-Time Computing and Robotics, RCAR 2025
Y2 - 1 June 2025 through 6 June 2025
ER -