arrow
Return

Statistical learning for OCR error correction

delete2018-11-01
delete17
PRE
AI
J
Jie Mei *
A
Aminul Islam
A
Abidalrahman Moh’d
Y
Yajing Wu
E
Evangelos Milios
DOI:10.1016/j.ipm.2018.06.001delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Modern OCR engines incorporate some form of error correction, typically based on dictionaries. However, there are still residual errors that decrease performance of natural language processing algorithms applied to OCR text. In this paper, we present a statistical learning model for post processing OCR errors, either in a fully automatic manner or followed by minimal user interaction to further reduce error rate. Our model employs web-scale corpora and integrates a rich set of linguistic features. Through an interdependent learning pipeline, our model produces and continuously refines the error detection and suggestion of candidate corrections. Evaluated on a historical biology book with complex error patterns, our model outperforms various baseline methods in the automatic mode and shows an even greater advantage when involving minimal user interaction. Quantitative analysis of each computational step further suggests that our proposed model is well-suited for handling volatile and complex OCR error patterns, which are beyond the capabilities of error correction incorporated in OCR engines.
Keywords:
OCR post-processing
OCR error
Error correction
Statistical learning
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

I
Information Processing and Management
IF:
6.9
Papers:
5.2K
Citations:
1.4W

Organization

U
university of louisiana lafayette
Scholars:
1.8K
Papers: 1.6K
Citations: 0
D
Dalhousie University
Scholars:
2.0W
Papers: 1.8W
Citations: 2.3W