arrow
Return

DiffuseDoc: Document geometric rectification via diffusion model

delete2025-10-24
delete0
PRE
AI
W
Wenfei Xiong
H
Huabing Zhou *
Y
Yanduo Zhang
T
Tao Lü
马佳义 cover
马佳义 (Jiayi Ma)
DOI:10.1016/j.cviu.2025.104554delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We propose DiffuseDoc, a comprehensive framework for geometric rectification of document images. To the best of our knowledge, DiffuseDoc is the first geometric document rectification model that adopts the diffusion model. • We leverage the rectification result as the control condition and derive the diffusion process with a learnable condition. By employing the diffusion prior, we effectively control the distribution of the rectifying result. • We contribute a dataset. The dataset comprises real-world document images captured by sensors, along with their corresponding scanned versions. • Extensive experiments show that our approach achieved state-of-the-art performance on benchmark datasets and our collected datasets, surpassing other methods in various evaluation metrics.

Journal

Computer Vision and Image Understanding cover
Computer Vision and Image Understanding
IF:
3.5
Papers:
428
Citations:
7.3K

Organization

W
wuhan institute of technology
Scholars:
1.0W
Papers: 6.5K
Citations: 11
W
wuhan university
Scholars:
8.0W
Papers: 5.8W
Citations: 70