返回
PSU: Particle Stacking Undersampling Method for Highly Imbalanced Big Data
DOI:10.1109/ACCESS.2020.3009753.png)
摘要
En 中文
Imbalanced classes are a common problem in machine learning, and the computational costs required for proper resampling increases with the data size. In this study, a simple and effective undersampling method, named particle stacking undersampling (PSU) was proposed. Compared with other competing undersampling methods, PSU can significantly reduce the computational costs, while minimizing information loss to prevent a prediction bias. The performance benchmark applied on 55 binary classification problems indicated that the proposed method not only achieved an enhanced classification performance over other well-known undersampling methods (random undersampling, NearMiss-1, NearMiss-2, cluster centroid, edited nearest neighbor, condensed nearest neighbor, and Tomek Links) but also provided a computational simplicity that can be scalable to large data. Moreover, an experiment verified that two propositions forming the basis of the PSU algorithm can also be applied to other undersampling methods to achieve methodological improvements.
Keyword:
Support vector machines
Training
Stacking
Computational efficiency
Classification algorithms
Licenses
Kernel
Data mining
imbalanced data
undersampling
big data
support vector machines
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Scanning mass spectrometer for quantitative reaction studies on catalytically active microstructures
Influences of Generator Parameters on Fault Current and Torque in a Large-Scale Superconducting Wind Generator发电机参数对大型超导风力发电机故障电流和转矩的影响
Arsenic removal from aqueous solutions by adsorption using novel MIL-53(Fe) as a highly efficient adsorbent使用新型MIL-53(Fe) 作为高效吸附剂通过吸附从水溶液中去除砷
RSC Advances
IF0

