arrow
Return

On Segment-Aware Monocular Depth Estimation Using Vision Transformers

delete2026-02-02
delete0
delete
OA
AI
V
Vasileios Arampatzakis *
G
George Pavlidis
N
Nikolaos Mitianoudis
N
Nikos Papamarkos
DOI:10.3390/info17020145delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Monocular Depth Estimation (MDE) infers per-pixel scene geometry from a single RGB image. Despite recent progress, global MDE models often blur depth discontinuities at object boundaries and fail to capture object-level structure. Segment-aware depth estimation addresses this limitation by exploiting semantic segmentation to decompose depth prediction into simpler, class-specific subproblems. In this work, we study semantic-aware MDE in a multi-branch design where each semantic class is handled by a lightweight Vision Transformer (ViT) branch that predicts dense depth for its class while suppressing interference from other regions. We further examine fusion strategies that merge the branch outputs into a single prediction: (i) a learnable cross-attention fusion module that predicts depth from the stack of per-class proposals and masks, and (ii) a parameter-free stitched summation that sums mask-gated outputs. The proposed architecture is simple, scalable, end-to-end trainable, and compatible with arbitrary transformer backbones. Experiments on Virtual KITTI 2, where ground-truth depth and semantic labels are available, show that segment-aware modeling produces sharper depth boundaries and improves standard error metrics compared to a single-branch baseline (AbsRel 0.243 -> 0.152; RMSE 11.952 -> 9.101). Finally, we find that the parameter-free summation matches, and in most cases improves upon, the accuracy of learned fusion while adding no computational overhead.
Keywords:
monocular depth estimation
semantic segmentation
vision transformers
segment-aware learning
depth fusion
Virtual KITTI 2
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

I
Information
IF:
2.9
Papers:
833
Citations:
9.8K

Organization

D
democritus university of thrace
Scholars:
1.3K
Papers: 521
Citations: 0