返回
An image-based automatic Arabic translation system
DOI:10.1016/j.patcog.2008.10.031.png)
摘要
En 中文
In this paper, we present a system that automatically translates Arabic text embedded in images into English. The system consists of three components: text detection from images, character recognition, and machine translation. We formulate the text detection as a binary classification problem and apply gradient boosting tree (GBT), support vector machine (SVM), and location-based prior knowledge to improve the F1 score of text detection from 78.95% to 87.05%. The detected text images are processed by off-the-shelf optical character recognition (OCR) software. We employ an error correction model to post-process the noisy OCR Output, and apply a bigram language model to reduce word segmentation errors. The translation module is tailored with compact data structure for hand-held devices. The experimental results show substantial improvements in both word recognition accuracy and translation quality. For instance, in the experiment of Arabic transparent font, the BLEU score increases from 18.70 to 33.47 with use of the error correction module. Published by Elsevier Ltd.
Keyword:
Text detection
Image classification
OCR
Error correction
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
Highly efficient electrocatalytic hydrogen production by nickel promoted molybdenum sulfide microspheres catalysts镍促进硫化钼微球催化剂的高效电催化制氢
RSC Advances
IF0
Video OCR: indexing digital news libraries by recognition of superimposed captions
MULTIMEDIA SYSTEMS
IF3.1
Physiology and tonotopic organization of auditory receptors in the cricketGryllus bimaculatus DeGeer

