返回
Probrank: a feature probability estimation-based framework for feature selection and ranking
DOI:10.1007/s12065-025-01118-7.png)
摘要
En 中文
特征选择是机器学习技术领域一个尚未得到充分探索的分支。它减少了不重要特征的数量以及训练集的规模。许多研究关注数值型特征。然而,特征也可以被分类(例如按颜色或类型),而这些类别通常与其他特征相关联。当我们把类别转换为数字时,这些关联关系就会丢失。因此,我们提出了一种基于特征概率估计的特征排序方法。本研究提出了一种从训练集中选择显著特征的方法,通过剔除不重要的特征来降低计算和存储复杂度。该方法适用于数值型和分类型数据。特征概率估计(FPE)通过移除不必要特征来提升系统的可靠性和执行速度。我们在七个不同数据集上执行了所提出的方法,并将其与包括PCA、K-best(卡方)、特征重要性、信息增益、互信息、相关性以及递归特征消除在内的流行特征选择技术进行了比较。实验结果表明,所提出的方法在速度方面优于现有特征选择方法,并且在精度上优于其中的许多方法;其性能也与最佳特征选择方法相当。
Keyword:
Feature selection
Feature probability estimation
Categorical data
Machine learning
Classification techniques
期刊
IF:
2.6
论文数:
120
被引数:
2.0K
机构
引用论文
SCADA intrusion detection scheme exploiting the fusion of modified decision tree and Chi-square feature selection
INTERNET OF THINGS
IF7.6
Optimizing text classification accuracy: a hybrid strategy incorporating enhanced NSGA-II and XGBoost techniques for feature selectionMosa, M.A. 优化文本分类精度:一种融合增强型NSGA-II和XGBoost技术进行特征选择的混合策略。Prog. Artif. Intell. 2025, 14, 275–299. [Google Scholar] [CrossRef]
A probability estimation-based feature reduction and Bayesian rough set approach for intrusion detection in mobile ad-hoc network
APPLIED INTELLIGENCE
IF3.5

