arrow
Return

Transformer-Based Reference Frame Synthesis for VVC Inter-Coding

delete2026-02-16
delete0
PRE
AI
Q
Qipu Qin
C
Cheolkon Jung
DOI:10.1109/TCSVT.2026.3664925delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
It has been demonstrated that deep learning generates high-quality reference frames and achieves considerable gains in inter coding. Existing neural network-based reference frame generation (NN-RFG) methods are mainly built on convolutional neural networks (CNNs) with limited receptive fields and are applied to the Versatile Video Coding (VVC). In this paper, we propose Transformer-based reference frame synthesis for VVC inter coding, named TRFS. We introduce Transformer into NN-RFG to capture global contextual information and construct multi-scale feature pyramids for accurate optical flow estimation. TRFS operates between the decoded picture buffer (DPB) and reference picture lists (RPL) in VVC, generating new reference frames through spatiotemporal compensation between previously reconstructed frames. First, we present a hierarchical feature extractor based on parameter-efficient Transformer to capture global contextual information with different resolutions and construct multi-scale feature pyramids. Second, we design a weight-sharing optical flow estimator consisting of residual blocks and Transformers to progressively refine the bidirectional or unidirectional optical flows in a coarse-to-fine manner. Third, we employ a U-net frame enhancer equipped with a ConvNeXt variant to learn residuals from the input and warped frames while removing blurring distortion caused by backward warping. To train the TRFS network, we utilize a two-stage incremental learning strategy based on quantization parameter (QP)-distance to address the imbalanced QP gap between the compressed input and its label, enhancing its learning capability. Various experiments demonstrate that TRFS achieves average Bjøntegaard Delta rate (BD-rate) gains of {RA: 6.61%, 13.86%, 13.30%} and {LB: 6.10%, 13.92%, 13.14%} for {Y, U, V} components over VTM-11.0_NNVC-10.0 with Neural Network-Based Video Coding (NNVC)-tools disabled. Furthermore, when evaluated against VTM-11.0_NNVC-10.0 with NNVC-tools enabled, TRFS delivers average BD-rate gains of {RA: 4.43% and LB: 3.40%} for the luma component.
Keywords:
Video coding
inter coding
inter prediction
reference frame generation
VVC
transformer
deep learning

Journal

IEEE Transactions on Circuits and Systems for Video Technology cover
IEEE Transactions on Circuits and Systems for Video Technology
IF:
11.1
Papers:
624
Citations:
3.1W

Organization

X
xidian university
Scholars:
6.5K
Papers: 2.2K
Citations: 0
Cited Papers

Cited Papers

No cited papers available