arrow
返回

Conditional Data Synthesis Augmentation

delete2026-01-21
delete0
PRE
AI
X
Xinyu Tian
X
Xiaotong Shen *
DOI:10.1080/01621459.2025.2586772delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Reliable machine learning and statistical analysis rely on diverse, well-distributed training data. However, real-world datasets are often limited in size and exhibit underrepresentation across key subpopulations, leading to biased predictions and reduced performance, particularly in supervised tasks such as classification. To address these challenges, we propose Conditional Data Synthesis Augmentation (CoDSA), a novel framework that leverages generative models, such as diffusion models, to synthesize high-fidelity data for improving model performance across multimodal domains, including tabular, textual, and image data. CoDSA generates synthetic samples that faithfully capture the conditional distributions of the original data, with a focus on under-sampled or high-interest regions. Through transfer learning, CoDSA fine-tunes pre-trained generative models to enhance the realism of synthetic data and increase sample density in sparse areas. This process preserves inter-modal relationships, mitigates data imbalance, improves domain adaptation, and boosts generalization. We also introduce a theoretical framework that quantifies the statistical accuracy improvements enabled by CoDSA as a function of synthetic sample volume and targeted region allocation, providing formal guarantees of its effectiveness. Extensive experiments demonstrate that CoDSA consistently outperforms non-adaptive augmentation strategies and state-of-the-art baselines in both supervised and unsupervised settings. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
Keyword:
Data augmentation
Generative models
Multimodality
Natural language processing
Transfer learning
Unstructured data

期刊

J
Journal of the American Statistical Association
IF:
3
论文数:
5.2K
被引数:
4.8W

机构

U
university of minnesota
学者数:
3.8K
论文数: 1.7K
被引数: 0
引用论文

引用论文

Universal inference
err2020-07-06
err75
errOAAI
errWasserman, Larry; Ramdas, Aaditya; Balakrishnan, Sivaraman
err分享
err收藏
HIGH-DIMENSIONAL VARIABLE SELECTION高维变量选择
err2009-10-01
err464
errOAAI
errWasserman, Larry; Roeder, Kathryn
err分享
err收藏
Random forests随机森林
err2001-01-01
err3.1W
errOAAI
errBreiman, L
err分享
err收藏
Convergence Rate of Sieve Estimates
err1994-06-01
err0
errOAAI
errXiaotong Shen; Wing Hung Wong
err分享
err收藏
学者 查看更多内容