Return
Saliency-guided video coding via recurrent learning and perceptual quality assessment
T
H
DOI:10.1016/j.image.2026.117536.png)
Abstract
En 中文
With the rapid growth of video streaming and conferencing, efficient video compression has become essential. This paper introduces a novel perceptually-optimized video compression framework, Saliency-guided Recurrent Learning Video Compression (SRLVC), which allocates bitrate to preserve high reconstruction quality in visually important regions while achieving efficient compression. The core innovation is a saliency-aware attention map that dynamically distinguishes between salient and non-salient areas, approximating the human visual sensitivity curve. By integrating this attention modulation mechanism with specialized motion and residual recurrent auto-encoders, SRLVC balances compression efficiency and perceptual quality. Another key contribution is the Saliency-based Peak Signal-to-Noise Ratio (SPSNR), a new perceptual quality assessment metric that exhibits strong correlation with human perception of video quality, as validated against Mean Opinion Scores (MOS). SPSNR provides a more perceptually relevant standard for evaluating video compression performance. Extensive experiments demonstrate SRLVC's superior rate-distortion performance, particularly in preserving the fidelity of salient regions even at low bitrates. By using visual saliency to guide both bitrate allocation and quality assessment, SRLVC achieves an effective balance between compression efficiency and perceived visual quality.
Keywords:
Recurrent learning
Learned video coding
Perceptual quality assessment
Saliency processing
Journal
S
IF:
2.7
Papers:
18
Citations:
0
