Return
DM-VSR: Depth-Aware Diffusion Models With Adaptive Modulation for Video Super-Resolution
L
Y
Y
Z
J
Y
DOI:10.1109/TBC.2025.3637713.png)
Abstract
En 中文
Video Super-Resolution (VSR) is essential for enhancing the quality of low-resolution (LR) videos in practical applications. Recent studies have explored diffusion models (DMs) for VSR due to their ability to generate realistic details. However, existing methods overlook spatiotemporal object-scale variations and dynamic control demands during denoising. This leads to visual distortion and quality degradation, severely limiting its practical applications. To address these limitations, we propose <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">DM-VSR</b>, a novel DM-based framework that incorporates depth-aware guidance and adaptive modulation for precise content reconstruction. Specifically, a Depth-aware Multimodal Fusion (DMF) module integrates depth maps, LR inputs, and flow-warped frames to provide unified depth-aware guidance. A Timestep Adaptive Modulation (TAM) module dynamically adjusts the control feature injection according to demand at each denoising step. Additionally, a Dynamic Consistency Loss (DCL) is introduced to align training objectives with the evolving semantic focus. Extensive experiments on the REDS4 and Vid4 benchmarks demonstrate that DM-VSR achieves competitive performance, surpassing state-of-the-art methods in both perceptual quality and temporal consistency. Moreover, DM-VSR generates more visually realistic results, emphasizing its effectiveness in real-world applications. The code will be released at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/aigcvsr/DM-VSR</uri>
Keywords:
Video super-resolution
diffusion models
depth map
perceptual quality
adaptive modulation
multimodal fusion
Journal
IF:
4.8
Papers:
2.1K
Citations:
3.0K
