Return
Boosting monocular depth estimation with semantic diffusion guided convolutional transformer model
DOI:10.1016/j.neucom.2025.132113.png)
Abstract
En 中文
• A dual-resolution estimation method to separate smooth base depth from residual detail depth is developed. Smooth base depth estimation focuses solely on wide-area consistency of depths while residual detail depth estimation addresses local regression of boundary depth. The boundary depth is, in fact, an exceedingly crucial component of the residual depth. It serves as the primary component that remains in the residual depth following the smoothing of the depth. For different branches, a contextual information encoder with varying ranges is designed based on convolution and transformer characteristics. • A semantically guided true boundary diffusion module is proposed. It adopts the approximate form of the spatial diffusion model and utilizes multi-scale contextual semantic information to guide depth diffusion at the anisotropic boundary, improving local boundary depth estimation. • Extensive experiments confirm that the proposed method performs well on public indoor and outdoor datasets, with some indicators equaling or surpassing state-of-the-art methods, particularly in terms of visual effects.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

