Return
Efficient Dataset Distillation via Generative Pruning
Y
M
G
K
S
D
DOI:10.1109/tbdata.2026.3668524.png)
Abstract
En 中文
Dataset distillation (DD) has demonstrated the promise of synthesizing smaller datasets that enable competitive performance. While most of DD methods operate in pixel space, they often suffer from poor scalability on high-resolution datasets. Recent works have thus shifted towards parameterizing synthetic data using deep generative priors. However, these approaches apply latent updates to all generator layers, leading to substantial computational overhead. To address this limitation, we propose Generative Lightweight Distillation (GLiD), a unified framework that jointly compresses the generator and accelerates latent optimization for efficient distillation. GLiD introduces two key components: (1) We introduce a Classification–Diversity Sensitivity Pruning mechanism that quantifies both discriminative utility and semantic diversity of each output channel to guide structural pruning and selective layer-wise optimization. (2) We also present a Layer-Adaptive Scheduling strategy that dynamically allocates latent update steps across generator stages based on convergence behavior. Extensive experiments on CIFAR-10 and ImageNet-1K and its subsets demonstrate that our GLiD achieves up to 10× acceleration across different datasets, while maintaining performance competitive with state-of-the-art methods.
Keywords:
Dataset distillation
pruning
generative model
Journal
I
IF:
5.7
Papers:
834
Citations:
3.0K
