arrow
Return

Feature Selection for High Dimensional Data Using Weighted K-Nearest Neighbors and Genetic Algorithm

delete2020-01-01
delete22
delete
OA
AI
S
Shuangjie Li
K
Kaixiang Zhang
Q
Qianru Chen
S
Shuqin Wang *
S
Shaoqiang Zhang
DOI:10.1109/ACCESS.2020.3012768delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Too many input features in applications may lead to over-fitting and reduce the performance of the learning algorithm. Moreover, in most cases, each feature containing different information content has different effects on the prediction target. Therefore, a feature selection method for calculating the importance of each feature, called WKNNGAFS, is proposed in this paper. In this method, the genetic algorithm (GA) is adopted to search the optimal weight vector, the value of the i th component of which corresponds to the contribution degree of the i th feature to the classification from a global perspective. Besides, weighted K-nearest neighbors algorithm (WKNN), which takes both the different contributions of nearest neighbors and the different classification ability of each feature into account, is used to determine the target label. To evaluate the effectiveness of the proposed method, nine existing feature selection methods are compared with it on 13 real datasets, including 6 high dimensional microarray datasets. Experimental results demonstrate the method is more effective and can improve classification performance.
Keywords:
Feature selection
weighted K-nearest neighbors
genetic algorithm
real coding
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

T
Tianjin Normal University
Scholars:
4.6K
Papers: 3.2K
Citations: 4.2K