arrow
Return

Partition-based differentially private synthetic data generation

delete2025-09-01
delete0
PRE
AI
M
Meifan Zhang
D
Deng, Dihang
L
Lihua Yin *
DOI:10.1016/j.ins.2025.122675delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Private synthetic data sharing is beneficial as it better retains the distribution and nuances of the original data compared to summary statistics such as means and frequencies. Current state-of-the-art methods follow a select-measure-generate paradigm, but measuring large-domain marginals often leads to significant errors, and managing the privacy budget poses challenges. Our partition-based approach addresses these issues, effectively reducing errors and improving the quality of synthetic data, even with a limited privacy budget. Experimental results show that our method outperforms existing approaches, yielding synthetic data with enhanced quality and utility, making it a preferred option for private data sharing.
Keywords:
Differential privacy
Synthetic data generation
Partition

Journal

Information Sciences cover
Information Sciences
IF:
6.8
Papers:
540
Citations:
6.2W

Organization

No organization information available