arrow
Return

Document retrieval tolerating character recognition errors - Evaluation and application

delete1997-08-01
delete17
PRE
AI
T
Tao Hu
H
Hiromichi Fujisawa
DOI:10.1016/S0031-3203(96)00155-0delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper presents two methods of combining character recognition with techniques for retrieving Japanese documents and also shows how these methods can be applied to textual image retrieval. Both retrieval methods are tolerant of errors that occur during the character recognition process. The basic idea is to utilize the characteristics of recognition errors. One uses a confusion matrix to generate ''equivalent'' query strings that should match erroneously recognized text. The other one searches a ''non-deterministic text'' that contains multiple candidates for ambiguous recognition results. Simulation experiments have shown that both methods can effectively combine character recognition with retrieval techniques. (C) 1997 Pattern Recognition Society. Published by Elsevier Science Ltd.
Keywords:
Japanese
character recognition
document retrieval
recognition error
confusion matrix
extended query-term method
non-deterministic text
multiple-candidate method
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

No organization information available
Cited Papers

Cited Papers

No cited papers available