arrow
Return

Enhanced Video Super-Resolution Method Using U-Net-Based Spatial Transformer Module

delete2026-02-17
delete0
PRE
AI
Y
Yooho Lee
S
S. M. Cho
D
Dongsan Jun
DOI:10.1109/TCSVT.2026.3665528delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Video super-resolution (VSR) aims to reconstruct high-resolution (HR) videos from low-resolution (LR) inputs by utilizing spatio-temporal correlations across consecutive LR frames. While recent advances in deep learning, particularly transformer-based architecture, have substantially improved VSR performance, maintaining spatio-temporal coherence remains a critical challenge. To address this issue, we propose a novel U-Net-based spatial transformer module (USTM) that can be seamlessly integrated into the reconstruction stage of existing VSR frameworks. The proposed USTM combines a spatial transformer with a U-Net structure to extract fine-grained spatio-temporal features to enhance the reconstruction of complex motions and texture patterns. Extensive ablation studies were conducted to identify the optimal configuration and verify the contribution of each USTM component. For performance evaluations, USTM was incorporated into multiple representative VSR methods. Experimental results demonstrate that integrating USTM consistently enhances PSNR and SSIM scores on benchmark datasets, including REDS, Vimeo-90K, and Vid4. Furthermore, visual comparisons highlight the superiority of the proposed method, particularly in high-frequency regions, compared to baseline VSR methods without USTM.
Keywords:
Video super-resolution
deep neural network
convolution neural network
U-Net
spatial transformer

Journal

IEEE Transactions on Circuits and Systems for Video Technology cover
IEEE Transactions on Circuits and Systems for Video Technology
IF:
11.1
Papers:
829
Citations:
3.1W

Organization

D
Dong-a University
Scholars:
411
Papers: 210
Citations: 0
E
Cited Papers

Cited Papers

No cited papers available