Return
From Structure to Detail: A Conditional Diffusion Framework for Extremely Low-Bitrate Image Compression
DOI:10.1016/j.sigpro.2025.110480.png)
Abstract
En 中文
Although existing diffusion-based methods produce visually rich textures at extremely low bitrates, they often sacrifice structural fidelity, resulting in significant deviations from the original image. To address this fundamental trade-off, we propose Fidelity-Perception Diffusion-based Image Compression (FPD-IC), a two-stage conditional diffusion framework that explicitly separates structure reconstruction and detail restoration. In Stage I, we use a VAE-based compressor to recover structurally faithful conditional images from highly compact bitstreams. In Stage II, a diffusion model, guided by the output from Stage I, generates visually rich details. This conditional approach allows the diffusion model to focus exclusively on perceptual enhancement while preserving the overall structure established in Stage I. Additionally, we introduce a lightweight Fidelity-Perception Tuner Module (FPTM) to combine the outputs of both stages, enabling controllable trade-offs between fidelity and perceptual quality. Extensive experiments on the Kodak and Tecnick datasets demonstrate the effectiveness and robustness of FPD-IC. On the Tecnick dataset, FPD-IC outperforms state-of-the-art diffusion-based methods by 2.24–3.57 dB in PSNR at bitrates below 0.06 bpp, while also achieving superior perceptual quality. Furthermore, FPD-IC shows strong robustness to input noise, consistently maintaining high fidelity and perceptual quality under Gaussian perturbations. The code will be released at https://github.com/mlkk518/FPD-IC .
Journal
IF:
3.6
Papers:
9.9K
Citations:
1.7W
Organization
No organization information available

