arrow
Return

Optimization of Skewed Data Using Sampling-Based Preprocessing Approach

delete2020-07-16
delete23
delete
OA
AI
S
Sushruta Mishra *
P
Pradeep Kumar Mallick
L
Lambodar Jena
G
Gyoo‐Soo Chae
DOI:10.3389/fpubh.2020.00274delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
In the past few years, classification has undergone some major evolution. With a constant surge of the amount of data gathered from different sources, efficient processing and analysis of data is becoming difficult. Due to the uneven distribution of data among classes, data classification with machine-learning techniques has become more tedious. While most algorithms focus on major data samples, they ignore the minor class data. Thus, the data-skewing issue is one of the critical problems that need attention of researchers. The paper stresses upon data preprocessing using sampling techniques to overcome the data-skewing problem. Here, three different sampling techniques such as Resampling, SpreadSubSampling, and SMOTE are implemented to reduce this uneven data distribution issue and classified with the K-nearest neighbor algorithm. The performance of classification is evaluated with various performance metrics to determine the efficiency of classification.
Keywords:
data skewing problem
machine learning
best first search
KNN algorithm
SMOTE
SpreadSubSampling
F-score
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Frontiers in Public Health cover
Frontiers in Public Health
IF:
3.4
Papers:
2.5W
Citations:
5.7W

Organization

Baekseok University cover
Baekseok University
Scholars:
100
Papers: 118
Citations: 51