arrow
Return

Pixel-Perfect Visual Geometry Estimation

delete2026-08-25
delete0
PRE
AI
G
Gangwei Xu
H
Haotong Lin
H
Hongcheng Luo
H
Haiyang Sun
B
Bing Wang
G
Guang Chen
S
Sida Peng
H
Hangjun Ye
杨新艳 cover
杨新艳 (Xin Yang)
DOI:10.1109/tpami.2026.3727248delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recovering clean and accurate geometry from images is essential for robotics and augmented reality. However, existing geometry foundation models still suffer severely from flying pixels and the loss of fine details. In this paper, we present pixel-perfect visual geometry models that can predict high-quality, flying-pixel-free point clouds by leveraging generative modeling in the pixel space. We first introduce Pixel-Perfect Depth (PPD), a monocular depth foundation model built upon pixel-space diffusion transformers (DiT). To address the high computational complexity associated with pixel-space diffusion, we propose two key designs: 1) Semantics-Prompted DiT, which incorporates semantic representations from vision foundation models to prompt the diffusion process, preserving global semantics while enhancing fine-grained visual details; and 2) Cascade DiT architecture that progressively increases the number of image tokens, improving both efficiency and accuracy. To further extend PPD to video (PPVD), we introduce a new Semantics-Consistent DiT, which extracts temporally consistent semantics from a multi-view geometry foundation model. We then perform reference-guided token propagation within the DiT to maintain temporal coherence with minimal computational and memory overhead. Our models achieve the best performance among all generative monocular and video depth estimation models and produce significantly cleaner point clouds than all other models. Code is available at https://github.com/gangweix/pixel-perfect-depth.
Keywords:
Modeling
Pixel
Videos
Depth measurement
Semantics
Geometry
Visualization
Diffusion models
Measurement
Training

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
1.0K
Citations:
9.8W

Organization

H
Huazhong University of Science and Technology
Scholars:
4.0K
Papers: 1.0K
Citations: 0
X
xiaomi ev
Scholars:
16
Papers: 6
Citations: 0
Z
zhejiang university
Scholars:
5.7K
Papers: 1.6K
Citations: 0
researcher View more organizations
Cited Papers

Cited Papers

No cited papers available