Return
Diffusion GAN-Based Oversampling for Imbalanced Tabular Data
DOI:10.1109/TKDE.2025.3639433.png)
Abstract
En 中文
Imbalanced class distribution disrupts the training of a classifier, resulting in biases favoring majority classes. Data oversampling is a common strategy to tackle this issue. However, traditional methods may generate incorrect and unnecessary instances when facing complex data challenges, such as class overlap, small disjuncts, and noise samples. Therefore, there is a need for an oversampling method that can accurately characterize the data distribution. This paper introduces a novel deep generative oversampling approach for balancing the imbalanced tabular data by leveraging diffusion models and Generative Adversarial Networks (GANs). The model comprises a generator constructed from diffusion models and a discriminator with a Noise-Sensitive Auxiliary Classifier (NSAC) and is trained through an adversarial process. The synergy of these two models enhances stability and sample quality compared to GANs, with faster sampling speed and better conditional generating ability than diffusion models. In experimental validation across 22 real-world datasets, our method consistently outperforms six counterparts regarding Accuracy, F1-score, and MCC for binary and multi-class scenarios. Notably, our approach enhances classifier accuracy for minority classes while maintaining a high level for the majority class, a facet often compromised by other algorithms.
Keywords:
Imbalanced data
tabular data
diffusion models
generative adversarial networks
oversampling
Journal
IF:
10.4
Papers:
6.7K
Citations:
3.2W

