arrow
Return

UniLayDet: Simple Multi-dataset Document Layout Analysis

delete2026-01-01
delete0
PRE
AI
P
Prasidh Srikumar *
A
Ajoy Mondal
C
C. V. Jawahar
DOI:10.1007/978-3-032-04614-7_3delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Information extraction from documents has become increasingly popular due to the rise of large language models (LLMs) and reaugmented generation (RAG) models. Document Layout Analysis (DLA) is a fundamental task in document AI, playing a crucial role in identifying semantically related elements within a document-a key step toward effective information extraction. Modern document layout analysis algorithms benefit from large-scale annotated datasets but suffer significant performance drops when tested across different datasets, limiting the generalization of models trained on a single source. To address this, we utilize a multi-dataset training approach for a Universal Layout Detection (UniLayDet) model utilizing a shared detection architecture with dataset-specific outputs and unifying the label space post-training through an automatic merging process. UniLayDet significantly improves generalization across datasets compared to models trained individually and also achieves competitive in-domain performance, notably attaining a mAP (@IoU[0.5-0.95]) of 68.9% on M(6)Doc (partitioned setting), close to the existing SOTA of 69.9%, showing that our model is simple yet effective for the task. Code and models are available at github.com/Mobius1D/UniLayDet.
Keywords:
Document layout analysis
multi-dataset training
universal layout detector
object detection

Journal

D
DOCUMENT ANALYSIS AND RECOGNITION-ICDAR 2025, PT I
IF:
0
Papers:
23
Citations:
0

Organization