Return
Ensemble feature selection using distance-based supervised and unsupervised methods in binary classification
DOI:10.1016/j.eswa.2022.116794.png)
Abstract
En 中文
Feature selection refers to the problem of finding the optimal subset of features by removing irrelevantand redundant features to improve classification accuracy. The determination of the most effective distancemeasures to evaluate the relevance and redundancy of features has not been investigated precisely todate. Moreover, the relation between relevancy and redundancy is still uncertain. This paper presents anovel relevancy-redundancy measurement based on distance applying the idea of the mRMR criteria to anunsupervised method. In addition, a supervised method is proposed, in which the features are ranked in termsof the distance between each pair of samples in different classes of the feature vector. Then an ensemble ofthe proposed supervised and unsupervised methods is applied to choose the most relevant features subset.This study investigates and compares the effects of 24 distance measures selected from five major families ofdistance functions on the performance of the proposed feature selection methods. The highest-ranked features are selected using an empirically achieved threshold. To evaluate the selected features, three classifiers, i.e.,Decision Tree, Support Vector Machine and Naive Bayes were applied to biomedical datasets representingbinary problems from the UCI data repository. The experimental results demonstrate the superiority of the proposed methods over the state-of-the-art and also classical feature selection ones in terms of improvingstability, classification accuracy, Recall (Sensitivity), Precision, F-measure, and Specificity
Keywords:
Feature selection
Supervised feature selection
Unsupervised feature selection
Distance measure
Filtering
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W

