arrow
Return

Facial video semantic coding for semantic communication

delete2025-06-01
delete0
PRE
AI
D
Du Qiyuan
Y
Yiping Duan
陶肖明 (Xiaoming Tao)
DOI:10.23919/JCC.ja.2023-0606delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multimedia semantic communication has been receiving increasing attention due to its significant enhancement of communication efficiency. Semantic coding, which is oriented towards extracting and encoding the key semantics of video for transmission, is a key aspect in the framework of multimedia semantic communication. In this paper, we propose a facial video semantic coding method with low bitrate based on the temporal continuity of video semantics. At the sender's end, we selectively transmit facial keypoints and deformation information, allocating distinct bitrates to different keypoints across frames. Compressive techniques involving sampling and quantization are employed to reduce the bitrate while retaining facial key semantic information. At the receiver's end, a GAN-based generative network is utilized for reconstruction, effectively mitigating block artifacts and buffering problems present in traditional codec algorithms under low bitrates. The performance of the proposed approach is validated on multiple datasets, such as VoxCeleb and TalkingHead-1kH, employing metrics such as LPIPS, DISTS, and AKD for assessment. Experimental results demonstrate significant advantages over traditional codec methods, achieving up to approximately 10-fold bitrate reduction in prolonged, stable head pose scenarios across diverse conversational video settings.
Keywords:
facial video
semantic coding
semantic communications
talking head
video compression

Journal

China Communications cover
China Communications
IF:
3.1
Papers:
1.8K
Citations:
5.0K

Organization