Return
A progressive scale difference learning network for video prediction
DOI:10.1016/j.neucom.2025.132500.png)
Abstract
En 中文
Accurate prediction of video sequences is crucial for numerous downstream tasks. Despite significant progress in this field, existing studies still face several challenges, such as reliance on additional auxiliary information (e.g.,optical flow, semantic map, etc.) and inaccurate predictions of moving object structures, which limit their practical applications to some extent. To address these limitations, this paper proposes a Progressive Scale Difference Learning Network (PSDLN), which relies solely on RGB images as input and adaptively enhances feature representations in motion regions. Specifically, PSDLN is composed of three key components: Spatial Aware Voxel Flow Module (SAVFM), Difference Learning Block (DLB), and Voxel Flow Interaction Block (VFIB).The SAVFM adopts a hierarchical architecture by stacking multiple modules with varying downsampling factors, enabling a coarse-to-fine refinement of voxel flow. The DLB is designed to capture voxel flow differences between adjacent scales, thereby enhancing the modeling of dynamic variations. The VFIB further strengthens temporal representation through an inter-frame interaction mechanism. By progressively extracting and integrating multi-scale information, PSDLN significantly improves prediction accuracy. Extensive experiments on multiple benchmark datasets demonstrate that PSDLN consistently outperforms state-of-the-art methods in both visual quality and prediction precision.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

