Return
Text identification for document image analysis using a neural network
DOI:10.1016/S0262-8856(98)00055-9.png)
Abstract
En 中文
A new bottom-up method is described that clusters the content of a mixed type document into text or non-text areas. The proposed approach is based on a new set of features combined with a self-organized neural network classifier. The set of features corresponds to the contents and the relationship of 3 x 3 masks, is selected by using a statistical reduction procedure, and provides texture information. Next, a Principal Components Analyzer (PCA) is applied, which results in a reduced number of 'effective' features. The final set of features is then utilized as input vector into a proper neural network to achieve the classification goal. The neural network classifier is based on a Kohonen Self Organized Feature Map (SOFM). Document blocks are classified as text, graphics, and halftones or to secondary subclasses corresponding to special cases of the primal classes. The proposed method can identify text regions included in graphics or even overlapped regions, that is, regions that cannot be separated with horizontal and vertical cuts. The performance of the method was extensively tested on a variety of documents with very promising results. (C) 1998 Elsevier Science B.V. All rights reserved.
Keywords:
block classification
document segmentation
page layout analysis
neural network classifiers
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
4.2
Papers:
4.1K
Citations:
6.7K
Organization
No organization information available
Cited Papers
New insight into the photoheterotrophic growth of the isocytrate lyase-lacking purple bacterium Rhodospirillum rubrum on acetate
Microbiology
IF0

