返回
Which standard classification algorithm has more stable performance for imbalanced network traffic data?
DOI:10.1007/s00500-023-09331-1.png)
摘要
En 中文
Most standard classification algorithms are difficult to effectively learn and predict from imbalanced network traffic data, which usually leads to lower classification accuracy. To analyze the influence of imbalanced network traffic data on the performance of standard classification algorithms, the imbalanced data augmentation algorithms are first designed to obtain the imbalanced network traffic data set with gradually varying Imbalance Ratio (IR) and belonging to the same distribution. Then, to obtain more objective classification result and simplify the evaluation process, the evaluation metric AFG is used to evaluate the classification performance of standard classification algorithms based on area under the receiver operating characteristic curve (AUC), F-measure and G-mean. Finally, based on AFG and coefficient of variation (CV), performance stability of standard classification algorithms on imbalanced network traffic data is obtained. Experiments of eight widely used standard classification algorithms on 25 different imbalanced network traffic data demonstrate that the classification performance of GNB, RF and DT is unstable, while BNB, KNN, LR, GBDT, and SVC are relatively stable and not susceptible to imbalanced data. Especially, the KNN has the most stable classification performance. Also, the results are statistically confirmed by Friedman and Nemenyi post hoc statistical tests.
Keyword:
Imbalanced network traffic data
Data augmentation algorithms
Standard classification algorithms
Stable classification performance
期刊
IF:
2.5
论文数:
1.0W
被引数:
2.1W
机构
引用论文
Oversampling technique based on fuzzy representativeness difference for classifying imbalanced data
APPLIED INTELLIGENCE
IF3.5
An ensemble imbalanced classification method based on model dynamic selection driven by data partition hybrid sampling数据划分混合采样驱动的模型动态选择集成不平衡分类方法
The impact of class imbalance in classification performance metrics based on the binary confusion matrix
PATTERN RECOGNITION
IF7.6
Arsenic removal from aqueous solutions by adsorption using novel MIL-53(Fe) as a highly efficient adsorbent使用新型MIL-53(Fe) 作为高效吸附剂通过吸附从水溶液中去除砷
RSC Advances
IF0
Preprocessed dynamic classifier ensemble selection for highly imbalanced drifted data streams
INFORMATION FUSION
IF15.5
Optimizing Weighted Extreme Learning Machines for imbalanced classification and application to credit card fraud detection
NEUROCOMPUTING
IF6.5
Study of the impact of resarnpling methods for contrast pattern based classifiers in imbalanced databases研究不平衡数据库中基于对比度模式的分类器的重新采样方法的影响
NEUROCOMPUTING
IF6.5

