Return
MRFGNet: Multiscale Reference Frame Generation Network for VVC Inter-Coding
DOI:10.1145/3750049.png)
Abstract
En 中文
Since the quality of the reference frames is critical for VVC inter-coding, the neural network (NN)-based reference frame generation aims to generate a better quality reference frame from the two decoded frames in the decoded picture buffer (DPB) and inserts it into the reference picture list (RPL) for inter-prediction. However, it is difficult to directly apply video frame interpolation or extrapolation to the reference frame generation because compression artifacts severely degrade the quality of the generated reference frame. In this article, we propose a multiscale reference frame generation network for VVC inter-coding, named MRFGNet. To reduce the compression artifacts, MRFGNet adopts the high-performance operation point (HOP) network, which has been released by JVET, as preprocessing for frame enhancement. Unlike the HOP in-loop filter that takes multiple inputs, MRFGNet only takes the reconstructed frame and quantization parameter (QP) map as input for the HOP network. Moreover, a frame generation network is proposed to conduct accurate motion estimation based on directional optical flow. Thus, in both RA and LDB configurations, MRFGNet has the same architecture to achieve both bidirectional and unidirectional predictions. MRFGNet estimates optical flow at multiple scales with the same dimension, thereby leveraging scale-independent bidirectional optical flow prediction. An optical flow warping and fusion module is designed to two cascaded U-nets for flow feature maps and get two frames, and finally, the two frames are fused to generate a reference frame. Furthermore, a novel training strategy based on QP distance is utilized to optimize MRFGNet by taking compressed data with higher quality as label for training. Experimental results show that MRFGNet achieves average BD rate changes of {−5.10% (Y), −12.14% (U), −11.51% (V)} and {−5.98% (Y), −15.32% (U), −14.22% (V)} over VTM_11.0-NNVC_4.0 anchor under RA and LDB configurations, respectively, as well as outperforms the state-of-the art method for reference frame generation, i.e., JVET-AD0160, by {0.78% (Y), 1.42% (U), 1.44% (V)} and {2.82% (Y), 4.92% (U), 6.62 %(V)} under RA and LDB configurations, respectively.
Journal
IF:
6
Papers:
2.0K
Citations:
5.4K
Organization
No organization information available

