arrow
返回

GAN-based multi-view video coding with spatio-temporal EPI reconstruction

delete2025-03-01
delete0
PRE
AI
C
Chengdong Lan
郝燕 封面图
郝燕 (Hao Yan)
C
Cheng Luo
T
Tiesong Zhao *
DOI:10.1016/j.image.2024.117242delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The introduction of multiple viewpoints in video scenes inevitably increases the bitrates required for storage and transmission. To reduce bitrates, researchers have developed methods to skip intermediate viewpoints during compression and delivery, and ultimately reconstruct them using Side Information (SInfo). Typically, depth maps are used to construct SInfo. However, these methods suffer from reconstruction inaccuracies and inherently high bitrates. In this paper, we propose a novel multi-view video coding method that leverages the image generation capabilities of Generative Adversarial Network (GAN) to improve the reconstruction accuracy of SInfo. Additionally, we consider incorporating information from adjacent temporal and spatial viewpoints to further reduce SInfo redundancy. At the encoder, we construct a spatio-temporal Epipolar Plane Image (EPI) and further utilize a convolutional network to extract the latent code of a GAN as SInfo. At the decoder, we combine the SInfo and adjacent viewpoints to reconstruct intermediate views using the GAN generator. Specifically, we establish a joint encoder constraint for reconstruction cost and SInfo entropy to achieve an optimal trade-off between reconstruction quality and bitrate overhead. Experiments demonstrate the significant improvement in Rate-Distortion (RD) performance compared to state-of-the-art methods.
Keyword:
Multi-view video coding
Generative adversarial network
Latent code learning
Epipolar plane image

期刊

S
Signal Processing and Image Communication
IF:
2.7
论文数:
2.8K
被引数:
4.2K

机构

F
fuzhou university
学者数:
3.3W
论文数: 2.1W
被引数: 31