arrow
Return

Supervised two-step feature extraction for structured representation of text data

delete2013-04-01
delete7
PRE
AI
O
Ondřej Háva *
M
Miroslav Skrbek
P
Pavel Kordík
DOI:10.1016/j.simpat.2012.11.003delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Training data matrix used for classification of text documents to multiple categories is characterized by large number of dimensions while the number of manually classified training documents is relatively small. Thus the suitable dimensionality reduction techniques are required to be able to develop the classifier. The article describes two-step supervised feature extraction method that takes advantage of projections of terms into document and category spaces. We propose several enhancements that make the method more efficient and faster than it was presented in our former paper. We also introduce the adjustment score that enables to correct defected targets or helps to identify improper training examples that bias extracted features. (C) 2012 Elsevier B.V. All rights reserved.
Keywords:
Document representation
Supervised feature extraction
Instance selection

Journal

Simulation Modelling Practice and Theory cover
Simulation Modelling Practice and Theory
IF:
4.6
Papers:
2.6K
Citations:
4.8K

Organization

C
czech technical university prague
Scholars:
6.5K
Papers: 5.3K
Citations: 3