Return
Dataset agnostic document object detection
DOI:10.1016/j.patcog.2023.109698.png)
Abstract
En 中文
Localizing document objects such as tables, figures, and equations is a primary step for extracting infor-mation from document images. We propose a novel end-to-end trainable deep network, termed Document Object Localization Network (doln et), for detecting various objects present in the document images. The proposed network is a multi-stage extension of Mask r-cnn with a dual backbone having deformable convolution for detecting document objects with high detection accuracy at a higher IoU threshold. We also empirically evaluate the proposed dolnet on the publicly available benchmark datasets. The pro-posed DOLNet achieves state-of-the-art performance for most of the bench-mark datasets under various existing experimental environments. Our solution has three important properties: (i) a single trained model dolnet & DDAG; that performs well across all the popular benchmark datasets, (ii) reports excellent performances across multiple, including with higher IoU thresholds, and (iii) consistently demonstrate the superior quantitative performance by fol-lowing the same protocol of the recent works for each of the benchmarks.& COPY; 2023 Elsevier Ltd. All rights reserved.
Keywords:
Document object detection
Table detection
Figure detection
Equation detection
Cascade Mask r-cnn
Deformable convolution
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

