Return
CLIP can understand depth
DOI:10.1016/j.patcog.2025.112475.png)
Abstract
En 中文
• We demonstrate that CLIP can be effectively generalized to monocular dense depth estimation using its existing prior knowledge, all without fine-tuning, validating our title “CLIP Can Understand Depth.” • We propose CLIP2Depth, a framework to correct CLIP prior for estimating depth by training a single non-human language embedding matrix called “mirror” with a lightweight deconvolutional Transformer decoder. • We show that a single query vector is sufficient to extract a single type of pixelwise features, eliminating the need for human-designed prompt sets or token binning. • Our model outperforms all previous CLIP-based depth estimation methods on NYU Depth v2 and KITTI. Furthermore, it performs competitively with taskspecific vision models while fully preserving the task-agnostic characteristic of the original CLIP.
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

