返回
Towards high dimensional instance selection: An evolutionary approach
DOI:10.1016/j.dss.2014.01.012.png)
摘要
En 中文
Data reduction is an important data pre-processing step in the KDD process. It can be approached by the application of some instance selection algorithms to filter out unrepresentative or noisy data from a given (training) dataset. However, the performance of instance selection over very high dimensional data has not yet been fully examined. In this paper, we introduce a novel efficient genetic algorithm (EGA), which fits biological evolution into the evolutionary process. In other words, after long-term evolution, individuals find the most efficient way to allocate resources and evolve. The experimental study is based on four very high dimensional datasets ranging from 200 to 18,236 dimensions. In addition, four state-of-the-art algorithms including IB3, DROP3, ICF, and GA are compared with EGA. The experimental results show that EGA allows the k-NN and SVM classifiers to provide the most comparable classification performance with the baseline classifiers without instance selection. Particularly, EGA outperforms the four algorithms in terms of average classification accuracy. Moreover, EGA can produce the largest reduction rates (the same as GA) and it requires relatively less computational time than the other four algorithms. (C) 2014 Elsevier B.V. All rights reserved.
Keyword:
Data reduction
Instance selection
Data mining
Machine learning
Genetic algorithms
High dimensional data
期刊
IF:
6.8
论文数:
3.8K
被引数:
1.5W
机构
引用论文
Electrical Conductivity of Reproductive Tissue for Detection of Estrus in Dairy Cows用于检测奶牛发情的繁殖组织电导率
Using evolutionary algorithms as instance selection for data reduction in KDD: An experimental study使用进化算法作为KDD中数据约简的实例选择: 一项实验研究

