返回
A distributed evolutionary based instance selection algorithm for big data using Apache Spark
DOI:10.1016/j.asoc.2024.111638.png)
摘要
En 中文
Instance selection is an important preprocessing technology in data mining and machine learning. In this paper, we proposed a novel evolutionary based instance selection algorithm for big data. First, we defined a coarse granularity chromosome structure to reduce the size of search space and costs of chromosome operations (recombination and mutation, etc.). Then a stratified evolution strategy was proposed to remove the hyper parameter in classic fitness function and achieve precise control over the reduction ratio of instances. Finally, a sampling-based fitness function was proposed to reduce the time complexity. Experimental results shown that our new algorithm is efficient to complete the instance selection task on data set with millions of instances in minutes-level. The 10-fold cross-validation also proved that the selection results on many datasets have high nearest neighbor classification accuracy.
Keyword:
Evolutionary algorithm
Instance selection
Apache Spark
Big Data
期刊
IF:
6.6
论文数:
1.4W
被引数:
4.8W
机构
暂无机构信息
引用论文
Spectral sensitivity of a novel photoreceptive system mediating entrainment of mammalian circadian rhythms
Nature
IF0
An Evolutionary Multiobjective Model and Instance Selection for Support Vector Machines With Pareto-Based Ensembles基于Pareto集成的支持向量机的进化多目标模型和实例选择
A high-dimensional feature selection method based on modified Gray Wolf Optimization基于改进灰狼优化的高维特征选择方法

