arrow
Return

Joint Semantic and Coded Generation for Conditional Latent Coding

delete2026-06-10
delete0
PRE
AI
S
Siqi Wu
W
Weiming Chen
Y
Yinda Chen
刘东 (Dong Liu)
K
K. C. Ho
何志海 (Zhihai He)
DOI:10.1109/tcsvt.2026.3702531delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep image compression and text-to-image generation represent two distinct paradigms in visual representation learning: one focuses on coded representations, while the other emphasizes semantic representations. This paper bridges this gap between these two by integrating image generation with compression. We propose Joint Semantic and Coded Generation for Conditional Latent Coding (JSCCL), which combines minimal coded features with semantic descriptions from vision-language models to guide a controllable rectified flow model. This joint guidance enables high-fidelity reference generation capturing both semantic coherence and structural details. Through iterative alignment in the latent domain, we synthesize conditional priors for efficient transformer-based encoding and decoding. Experimental results on Kodak, CLIC, and Tecnick datasets demonstrate up to 15.8% BD-rate improvement over VVC with only 0.01 bpp overhead. Unlike existing generative compression methods which approximate the inputs solely at the semantic level, our approach achieves both semantic and pixel-level precise reconstruction.
Keywords:
Image compression
generative models
rectified flow
conditional coding
semantic representation
deep learning

Journal

IEEE Transactions on Circuits and Systems for Video Technology cover
IEEE Transactions on Circuits and Systems for Video Technology
IF:
11.1
Papers:
624
Citations:
3.1W

Organization

U
university of missouri
Scholars:
1.3K
Papers: 607
Citations: 0
S
southern university of science and technology
Scholars:
4.6K
Papers: 1.7K
Citations: 0
U
University of Science and Technology of China
Scholars:
1.7W
Papers: 5.9K
Citations: 11.3W
researcher View more organizations