Return
Sequence-based undersampling: an algorithm for managing imbalanced datasets
DOI:10.7717/peerj-cs.3078.png)
Abstract
En 中文
Imbalanced datasets pose significant challenges in machine learning (ML), often leading to catastrophic failures in models that are insensitive to their statistical particularities. In regression using neural networks (NN), imbalance is particularly problematic due to the potential undersampling of extreme but important values. Traditional classification methods are ineffective with imbalanced datasets and fail to capture the structure of the problem. This article presents a novel approach to handling imbalanced data in general regression. The new preprocessing technique, called Sequence-Based Undersampling, leverages the spatial structure of the data to selectively remove overrepresented instances. The method is tested using quantitative precipitation estimates (QPE), a well-known case of imbalance distribution in earth physics. The technique demonstrates consistent improvements in model performance compared to existing methods. The results suggest that sequence-aware undersampling improves regression models and ML algorithms, providing a practical solution to a prevalent issue in data-driven research. This method can enhance current satellite precipitation algorithms, as satellite retrievals often exhibit a leptokurtic distribution with few cases of high rainfall rates, many low rates, and numerous no-rain occurrences-a paradigmatic case of an imbalanced dataset built sequentially through radar and radiometer measurements.
Keywords:
Imbalance
ML
IA
Regression
Earth sciences
Undersampling
Journal
IF:
2.5
Papers:
3.4K
Citations:
6.9K

