返回
A Solution to the High-Dimensional Classification Problem Using an Improved Hybrid Feature Selection Algorithm Guided by Interaction Information
DOI:10.1109/ACCESS.2020.3014825.png)
摘要
En 中文
This paper addresses the high-dimensional classification problem, which is very important in machine learning. When the number of features of the data is very high, the classification performance of a given classifier can degrade because there are not enough samples for training. One of the solutions to cope with this problem is to perform feature selection to reduce the number of features. We propose a new hybrid feature selection algorithm based on interaction information that improves upon the previous one. Our improved method employs interaction information to select candidate features to be added to the current feature subset. Cohen's d is used as the significance testing to decide whether a new feature is permanently added to the subset. We adopt new stopping criteria to allow intensive search. Our search method is efficient and is able to find excellent solutions. Experiments results on eleven high-dimensional data sets show that compared to other hybrid feature selection algorithms, our proposed algorithm provides high classification accuracy and requires a small number of features for classification.
Keyword:
Feature selection
high-dimensional data
hybrid algorithm
interaction information
sequential search
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Detecting thermophilic proteins through selecting amino acid and dipeptide composition features
AMINO ACIDS
IF2.4
Suboptimal branch and bound algorithms for feature subset selection: A comparative study特征子集选择的次优分支定界算法: 比较研究
Gene selection for cancer classification using support vector machines使用支持向量机进行癌症分类的基因选择
MACHINE LEARNING
IF2.9

