arrow
Return

ADOCRNet: A Deep Learning OCR for Arabic Documents Recognition

delete2024-01-01
delete5
delete
OA
AI
L
Lamia Mosbah *
I
Ikram Moalla
T
Tarek M. Hamdani
B
Bilel Neji
T
Taha Beyrouthy
A
Adel M. Alimi
DOI:10.1109/ACCESS.2024.3379530delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In recent years, Optical character recognition (OCR) has experienced a resurgence of interest especially for contemporary Arabic data. In fact, OCR development for printed and handwritten Arabic script is still a challenging task. These challenges are due to the specific characteristics of the Arabic script. In this work, we attempt to address these challenges by creating a deep learning OCR for Arabic document recognition called ADOCRNet. It is a novel deep learning framework whose architecture is built of layers of Convolutional Neural Networks (CNNs) and Bidirectional Long Short-Term Memory (BLSTM) trained using Connectionist Temporal Classification (CTC) algorithm. In order to assess the performance of our OCR, the proposed system is performed on two printed text datasets which are P-KHATT (text line images) and APTI (word images). It's also evaluated on a handwritten Arabic text dataset IFN/ENIT (word images). According to the practical tests, the conceived model achieves strength recognition rates on the three datasets. ADOCRNet reaches a Character Error Rate (CER) of 0.01% on the P-KHATT dataset, 0.03% on the APTI dataset and a Word Error Rate (WER) of 1.09% on the IFN/ENIT dataset, which significantly outperforms the outcomes of the current systems.
Keywords:
Arabic
document recognition
CNNs
CTC
deep learning
BLSTM
OCR

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.7W
Citations:
29.4W

Organization

U
universite de monastir
Scholars:
5.9K
Papers: 4.7K
Citations: 2
U
universite de sfax
Scholars:
8.9K
Papers: 7.7K
Citations: 5