arrow
返回

Multi-class imbalanced big data classification on Spark

delete2021-01-01
delete55
PRE
AI
W
William C. Sleeman
B
Bartosz Krawczyk *
DOI:10.1016/j.knosys.2020.106598delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Despite more than two decades of progress, learning from imbalanced data is still considered as one of the contemporary challenges in machine learning. This has been further complicated by the advent of the big data era, where popular algorithms dedicated to alleviating the class skew impact are no longer feasible due to the volume of datasets. Additionally, most of existing algorithms focus on binary imbalanced problems, where majority and minority classes are well-defined. Multi-class imbalanced data poses further challenges as the relationship between classes is much more complex and simple decomposition into a number of binary problems leads to a significant loss of information. In this paper, we propose the first compound framework for dealing with multi-class big data problems, addressing at the same time the existence of multiple classes and high volumes of data. We propose to analyze the instance-level difficulties in each class, leading to understanding what causes learning difficulties. We embed this information in popular resampling algorithms which allows for informative balancing of multiple classes. We propose an efficient implementation of the discussed algorithm on Apache Spark, including a novel version of SMOTE that overcomes spatial limitations in distributed environments of its predecessor. Extensive experimental study shows that using instance-level information significantly improves learning from multi-class imbalanced big data. Our framework can be downloaded from https://github.com/fsleeman/minority-type-imbalanced. (C) 2020 Elsevier B.V. All rights reserved.
Keyword:
Machine learning
Big data
Imbalanced data classification
Multi-class imbalance
Spark
MapReduce
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.2W
被引数:
4.5W

机构

V
Virginia Commonwealth University
学者数:
2.2W
论文数: 1.8W
被引数: 1.9W
引用论文

引用论文

Seed oils fromCitrus sinensis
err1963-12-01
err0
PREAI
errR. Hendrickson; J. W. Kesterson
err分享
err收藏
err分享
err收藏
err分享
err收藏
err分享
err收藏
Atypical 22q11.2 deletion in a patient with DGS/VCFS spectrum
err2008-05-01
err0
errOAAI
errSintia Iole Nogueira; April M. Hacker; Fernanda T.S. Bellucco; Denise M. Christofolini; Leslie Domenici Kulikowski; Mirlene C.S.P. Cernach; Beverly S. Emanuel; Maria Isabel Melaragno
err分享
err收藏
Mineral composition of fruit by-products evaluated by neutron activation analysis
err2013-01-13
err0
PREAI
errGabriela de Matuoka e Chiocchetti; Elisabete A. De Nadai Fernandes; Márcio Arruda Bacchi; Rogério Augusto Pazim; Silvana Regina Vicino Sarriés; Thaís Melega Tomé
err分享
err收藏
A study on combining dynamic selection and data preprocessing for imbalance learning
err2018-04-01
err89
PREAI
errRoy, Anandarup; Cruz, Rafael M. O.; Sabourin, Robert; Cavalcanti, George D. C.
err分享
err收藏
An Information Theory-Based Feature Selection Framework for Big Data Under Apache SparkApache Spark下基于信息论的大数据特征选择框架
err2018-09-01
err49
PREAI
errRamirez-Gallego, Sergio; Mourino-Talin, Hector; Martinez-Rego, David; Bolon-Canedo, Veronica; Manuel Benitez, Jose; Alonso-Betanzos, Amparo; Herrera, Francisco
err分享
err收藏
学者 查看更多内容