arrow
Return

CLIP can understand depth

delete2025-09-26
delete0
PRE
AI
S
Sohee Kim
J
Jisu Kang
D
Dunam Kim
S
Seokju Lee
DOI:10.1016/j.patcog.2025.112475delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We demonstrate that CLIP can be effectively generalized to monocular dense depth estimation using its existing prior knowledge, all without fine-tuning, validating our title “CLIP Can Understand Depth.” • We propose CLIP2Depth, a framework to correct CLIP prior for estimating depth by training a single non-human language embedding matrix called “mirror” with a lightweight deconvolutional Transformer decoder. • We show that a single query vector is sufficient to extract a single type of pixelwise features, eliminating the need for human-designed prompt sets or token binning. • Our model outperforms all previous CLIP-based depth estimation methods on NYU Depth v2 and KITTI. Furthermore, it performs competitively with taskspecific vision models while fully preserving the task-agnostic characteristic of the original CLIP.

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

K
Korea Institute of Energy Technology
Scholars:
94
Papers: 56
Citations: 0