arrow
返回

Dual-Conditioned Training to Exploit Pre-Trained Codebook-Based Generative Model in Image Compression

delete2024-01-01
delete0
delete
OA
AI
S
Shoma Iwai *
T
Tomo Miyazaki
S
Shinichiro Omachi
DOI:10.1109/ACCESS.2024.3522238delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Learned image compression (LIC) is increasingly gaining attention. To improve the perceptual quality of reconstructions, generative LIC has been studied, using generative models such as Generative Adversarial Networks (GANs). State-of-the-art generative LIC methods have achieved remarkable performance even in low bit rate settings. Unlike most approaches trained from scratch, we propose a generative LIC that utilizes a pre-trained codebook-based generative model, Vector-Quantized GAN (VQGAN). Specifically, our model is designed to exploit its powerful image-generation capabilities to enhance compression performance. Our approach reconstructs an image from a transmitted bitstream in two steps: (1) estimating VQGAN tokens and feeding them into the pre-trained VQGAN decoder, and (2) modifying the decoder's intermediate features to address artifacts and distortions. Our preliminary experiments reveal that the information allocation between (1) and (2) is pivotal for reconstruction quality. Moreover, we found that the ideal allocation varies based on the target bit rate. Motivated by these findings, we propose a novel Dual-Conditioned training. Through the training, the model learns to adjust the total bit rate and information allocation between (1) and (2) based on two conditional inputs. Subsequently, we explore the conditional inputs to achieve the optimal results for each target bit rate. This training strategy enables us to effectively exploit the generation capability of VQGAN across different bit rates. Our method, named Dual Conditioned VQGAN-based Image Compression (DC-VIC), outperforms state-of-the-art generative LIC methods in rate-distortion-perception performance.
Keyword:
Image coding
Decoding
Image reconstruction
Training
Bit rate
Entropy
Estimation
Resource management
Distortion
Transforms
Generative adversarial networks
image compression
VQGAN

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

T
tohoku university
学者数:
4.3W
论文数: 3.6W
被引数: 31
引用论文

引用论文

MoVam7, a Conserved SNARE Involved in Vacuole Assembly, Is Required for Growth, Endocytosis, ROS Accumulation, and Pathogenesis of Magnaporthe oryzae
err2011-01-24
err0
errOAAI
errXianying Dou; Qi Wang; Zhongqiang Qi; Wenwen Song; Wei Wang; Min Guo; Haifeng Zhang; Zhengguang Zhang; Ping Wang; Xiaobo Zheng
err分享
err收藏
err分享
err收藏
How Himalayan collision stems from subduction
err2021-04-28
err0
PREAI
errM. Soret; K.P. Larson; J. Cottle; A. Ali
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Ki-67 Proliferation Index
err2004-02-01
err0
PREAI
errMichael G. Alexandrakis; Freda H. Passam; Despina S. Kyriakou; Konstantina Dambaki; Maria Niniraki; Efstathios Stathopoulos
err分享
err收藏
Left ventricular noncompaction (LVNC) and low mitochondrial membrane potential are specific for Barth syndrome
err2013-01-30
err0
errOAAI
errAgnieszka Karkucinska‐Wieckowska; Joanna Trubicka; Bozena Werner; Katarzyna Kokoszynska; Magdalena Pajdowska; Maciej Pronicki; Elzbieta Czarnowska; Magdalena Lebiedzinska; Jolanta Sykut‐Cegielska; Lidia Ziolkowska; Weronika Jaron; Anna Dobrzanska; Elzbieta Ciara; Mariusz R. Wieckowski; Ewa Pronicka
err分享
err收藏
err分享
err收藏
学者 查看更多内容