Return
Joint Semantic and Coded Generation for Conditional Latent Coding
DOI:10.1109/tcsvt.2026.3702531.png)
Abstract
En 中文
Deep image compression and text-to-image generation represent two distinct paradigms in visual representation learning: one focuses on coded representations, while the other emphasizes semantic representations. This paper bridges this gap between these two by integrating image generation with compression. We propose Joint Semantic and Coded Generation for Conditional Latent Coding (JSCCL), which combines minimal coded features with semantic descriptions from vision-language models to guide a controllable rectified flow model. This joint guidance enables high-fidelity reference generation capturing both semantic coherence and structural details. Through iterative alignment in the latent domain, we synthesize conditional priors for efficient transformer-based encoding and decoding. Experimental results on Kodak, CLIC, and Tecnick datasets demonstrate up to 15.8% BD-rate improvement over VVC with only 0.01 bpp overhead. Unlike existing generative compression methods which approximate the inputs solely at the semantic level, our approach achieves both semantic and pixel-level precise reconstruction.
Keywords:
Image compression
generative models
rectified flow
conditional coding
semantic representation
deep learning
Journal
IF:
11.1
Papers:
624
Citations:
3.1W

