返回
Word spotting in historical documents using primitive codebook and dynamic programming
DOI:10.1016/j.imavis.2015.09.006.png)
摘要
En 中文
Word searching and indexing in historical document collections are a challenging problem because text characters are often touching or broken due to degradation or aging effects. In this paper, we present a novel approach towards word spotting using text line decomposition into character primitives and string matching. The text lines are initially separated by a segmentation process. Then each text line is described as sequences of primitive labels which correspond to single characters or parts of characters. These representative primitives are considered from a codebook of shapes generated from training pages taken from the collection. During indexation, the text lines are transcribed into strings of primitives in off-line stage and stored in files. For this purpose, an efficient indexation strategy using multi-label approach is used by a combination of two-level analysis of the primitives: coarse and fine levels. During retrieval, the query word image is encoded into strings of coarse and fine primitives chosen according to the codebook. Finally, a dynamic programming method based on approximate string matching is used to find similar primitive sequences in the text lines from the collection in runtime. We present the experimental evaluation on datasets of real life document images, gathered from historical books of different scripts. Experimental results show that the method is robust in searching text in noisy documents. (C) 2015 Elsevier B.V. All rights reserved.
Keyword:
Word spotting
Document indexing
Approximate string matching
Coarse-to-fine
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
4.2
论文数:
4.1K
被引数:
6.7K
机构
引用论文
A synthesised word approach to word retrieval in handwritten documents手写文档中单词检索的综合单词方法
PATTERN RECOGNITION
IF7.6

