Return
A Systematic Framework for Compressing Generative Diffusion Models for Resource-Constrained IoT Devices
DOI:10.1109/JIOT.2025.3613700.png)
Abstract
En 中文
Generative diffusion models deliver remarkable synthesis quality but remain impractical for resource-limited Internet of Things (IoT) devices due to their substantial computational and memory demands. To bridge this critical gap, we present a comprehensive, multistage optimization framework that systematically reduces model size while meticulously preserving generative fidelity. The framework integrates an efficient backbone architecture designed for inherent lightness, a sensitivity-guided fine-grained pruning strategy that strategically removes redundant parameters to achieve high sparsity, and a novel distribution-aware quantization algorithm based on Gaussian mixture models (GMMs) to compress weights and activations with minimal quality degradation. Extensive validation across multiple diffusion architectures [Denoising diffusion probabilistic models (DDPM), denoising diffusion implicit models (DDIM), and score-based generative model (SGM)] and diverse datasets demonstrates the framework’s strong generalizability, achieving up to 79% model sparsity while preserving generative fidelity. To showcase practical utility, we demonstrate that our framework produces a compressed model compatible with standard mobile deployment toolchains, realizing a significant reduction in the on-device memory footprint required for inference. This work offers a robust and generalizable methodology for enabling advanced generative artificial intelligence (AI) on a wide spectrum of edge and IoT platforms. The code is available at https://github.com/mitchell-cheng/compress_diffusion
Keywords:
Diffusion models
Gaussian mixture models (GMMs)
model compression
model pruning
model quantization
Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W

