arrow
Return

DiffuVolume: Diffusion Model for Volume based Stereo Matching

delete2025-02-01
delete0
PRE
AI
D
Dian Zheng
X
Xiao-Ming Wu
Z
Zuhao Liu
J
Jingke Meng
W
Wei‐Shi Zheng *
DOI:10.1007/s11263-025-02362-1delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Stereo matching is a significant part in many computer vision tasks and driving-based applications. Recently cost volume-based methods have achieved great success benefiting from the rich geometry information in paired images. However, the redundancy of cost volume also interferes with the model training and limits the performance. To construct a more precise cost volume, we pioneeringly apply the diffusion model to stereo matching. Our method, termed DiffuVolume, considers the diffusion model as a cost volume filter, which will recurrently remove the redundant information from the cost volume. Two main designs make our method not trivial. Firstly, to make the diffusion model more adaptive to stereo matching, we eschew the traditional manner of directly adding noise into the image but embed the diffusion model into a task-specific module. In this way, we outperform the traditional diffusion stereo matching method by 27%\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\%$$\end{document} EPE improvement and 7 times parameters reduction. Secondly, DiffuVolume can be easily embedded into any volume-based stereo matching network, boosting performance with only a slight increase in parameters (approximately 2%\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\%$$\end{document}). By adding the DiffuVolume into well-performed methods, we outperform all the published methods on Scene Flow, KITTI2012, KITTI2015 benchmarks and zero-shot generalization setting. It is worth mentioning that the proposed model ranks 1st on KITTI 2012 leader board, 2nd on KITTI 2015 leader board since 15, July 2023.
Keywords:
Stereo matching
Cost volume
Information filtering
Task-specific diffusion process

Journal

International Journal of Computer Vision cover
International Journal of Computer Vision
IF:
9.3
Papers:
3.9K
Citations:
2.8W

Organization

M
Minist Educ
Scholars:
1.4K
Papers: 566
Citations: 122
S
sun yat sen university
Scholars:
1.2W
Papers: 3.9K
Citations: 1.2K
Cited Papers

Cited Papers

U-Net: Convolutional Networks for Biomedical Image Segmentation
err2015-11-18
err0
PREAI
errOlaf Ronneberger; Philipp Fischer; Thomas Brox
errShare
errSave
errShare
errSave
A Multi-view Stereo Benchmark with High-Resolution Images and Multi-camera Videos
err2017-07-01
err0
PREAI
errThomas Schops; Johannes L. Schonberger; Silvano Galliani; Torsten Sattler; Konrad Schindler; Marc Pollefeys; Andreas Geiger
errShare
errSave
Adaptive Unimodal Cost Volume Filtering for Deep Stereo Matching
err2020-04-03
err0
errOAAI
errYoumin Zhang; Yimin Chen; Xiao Bai; Suihanjin Yu; Kun Yu; Zhiwei Li; Kuiyuan Yang
errShare
errSave
Group-Wise Correlation Stereo Network
err2019-06-01
err0
errOAAI
errXiaoyang Guo; Kai Yang; Wukui Yang; Xiaogang Wang; Hongsheng Li
errShare
errSave
researcher View more