返回
Display text segmentation after learning best-fitted OCR binarization parameters
DOI:10.1016/j.eswa.2011.09.162.png)
摘要
En 中文
In this paper text segmentation in generic displays is proposed through learning the best binarization values for a commercial optical character recognition (OCR) system. The commercial OCR is briefly introduced as well as the parameters that affect the binarization for improving the classification scores. The purpose of this work is to provide the capability to automatically evaluate standard textual display information, so that tasks that involve visual user verification can be performed without human intervention. The problem to be solved is to recognize text characters that appear on the display, as well as the color of the characters' foreground and background. The paper introduces how the thresholds are learnt through: (a) selecting lightness or hue component of a color input cell, (b) enhancing the bitmaps' quality, and (c) calculating the segmentation threshold range for this cell. Then, starting from the threshold ranges learnt at each display cell, the best threshold for each cell is gotten. The input and output data sets for testing the algorithms proposed are described, as well as the analysis of the results obtained. (C) 2011 Elsevier Ltd. All rights reserved.
Keyword:
Optical character recognition
Text segmentation
Binarization
Learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W
机构
引用论文
OCR binarization and image pre-processing for searching historical documents
PATTERN RECOGNITION
IF7.6
Low resolution, degraded document recognition using neural networks and hidden Markov models使用神经网络和隐马尔可夫模型进行低分辨率,降级的文档识别

