arrow
Return

From Structure to Detail: A Conditional Diffusion Framework for Extremely Low-Bitrate Image Compression

delete2026-01-08
delete0
PRE
AI
J
Junhui Li
Y
Yiyang Zou
侯兴松 cover
侯兴松 (Xingsong Hou) *
Y
Yutao Zhang
Z
Zhixuan Guo
DOI:10.1016/j.sigpro.2025.110480delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Although existing diffusion-based methods produce visually rich textures at extremely low bitrates, they often sacrifice structural fidelity, resulting in significant deviations from the original image. To address this fundamental trade-off, we propose Fidelity-Perception Diffusion-based Image Compression (FPD-IC), a two-stage conditional diffusion framework that explicitly separates structure reconstruction and detail restoration. In Stage I, we use a VAE-based compressor to recover structurally faithful conditional images from highly compact bitstreams. In Stage II, a diffusion model, guided by the output from Stage I, generates visually rich details. This conditional approach allows the diffusion model to focus exclusively on perceptual enhancement while preserving the overall structure established in Stage I. Additionally, we introduce a lightweight Fidelity-Perception Tuner Module (FPTM) to combine the outputs of both stages, enabling controllable trade-offs between fidelity and perceptual quality. Extensive experiments on the Kodak and Tecnick datasets demonstrate the effectiveness and robustness of FPD-IC. On the Tecnick dataset, FPD-IC outperforms state-of-the-art diffusion-based methods by 2.24–3.57 dB in PSNR at bitrates below 0.06 bpp, while also achieving superior perceptual quality. Furthermore, FPD-IC shows strong robustness to input noise, consistently maintaining high fidelity and perceptual quality under Gaussian perturbations. The code will be released at https://github.com/mlkk518/FPD-IC .

Journal

Signal Processing cover
Signal Processing
IF:
3.6
Papers:
9.9K
Citations:
1.7W

Organization

No organization information available