返回
Multi-Objective Feature Selection With Missing Data in Classification
DOI:10.1109/TETCI.2021.3074147.png)
摘要
En 中文
Feature selection (FS) is an important research topic in machine learning. Usually, FS is modelled as a bi-objective optimization problem whose objectives are: 1) classification accuracy; 2) number of features. One of the main issues in real-world applications is missing data. Databases with missing data are likely to be unreliable. Thus, FS performed on a data set missing some data is also unreliable. In order to directly control this issue plaguing the field, we propose in this study a novel modelling of FS: we include reliability as the third objective of the problem. In order to address the modified problem, we propose the application of the non-dominated sorting genetic algorithm-III (NSGA-III). We selected six incomplete data sets from the University of California Irvine (UCI) machine learning repository. We used the mean imputation method to deal with the missing data. In the experiments, k-nearest neighbors (K-NN) is used as the classifier to evaluate the feature subsets. Experimental results show that the proposed three-objective model coupled with NSGA-III efficiently addresses the FS problem for the six data sets included in this study.
Keyword:
Sorting
Statistics
Sociology
Optimization
Feature extraction
Genetic algorithms
Machine learning
Feature selection
Multi-objective
Optimization
NSGA-III
Missing data
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
I
IF:
6.5
论文数:
1.4K
被引数:
4.5K
机构
引用论文
Is Arenicola marina a suitable test organism to evaluate the bioaccumulation potential of Hg, PAHs and PCBs from dredged sediments?
Chemosphere
IF0
Variable-Length Particle Swarm Optimization for Feature Selection on High-Dimensional Classification高维分类特征选择的变长粒子群算法

