arrow
Return

Interactive image-to-video transfer learning

delete2025-11-25
delete0
PRE
AI
C
Cong Wu
T
Tianyang Xu
Z
Zhenhua Feng
X
Xiao‐Jun Wu
J
Josef Kittler
DOI:10.1016/j.neunet.2025.108354delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Transfer learning from image to video has become a widely adopted strategy in action recognition. Existing mainstream approaches typically fine-tune the entire network after initialising with pre-trained parameters, which inevitably leads to substantial training overhead. Recent efficiency-oriented methods alleviate this by freezing the backbone during optimisation; however, they tend to overly rely on powerful prior knowledge encoded in image models and often neglect how to effectively perform video reasoning with static backbones. In this paper, we propose an efficient image-to-video transfer learning framework, termed SDST (Static-Dynamic & Spatial-Temporal interaction framework), which explicitly enhances the interaction between static and dynamic cues as well as spatial and temporal domains. Specifically, we introduce the Motion Booster Module that extracts motion descriptors from intermediate features and integrates them via the progressive cross-attention mechanism to fuse static and dynamic representations. Furthermore, we propose a lightweight temporal modelling module, named Channel-Aware Multi-Scale Temporal Modelling, characterised by channel awareness, various temporal scales modelling, as well as feature interactions. By enriching the frozen image backbone through these components, our approach effectively bridges the gap between static visual representations and video-based action recognition tasks. Extensive experiments on Something-Something V1&V2, Diving-48, and Kinetics-400 demonstrate that our method not only surpasses State-of-the-Art efficient transfer learning techniques but also outperforms several fully fine-tuned approaches, all without additional bells and whistles. Moreover, we validate the transferability of the proposed SDST framework to vision-language pre-trained models like CLIP, achieving further gains in performance. These results highlight the potential of our method as a general and scalable solution for efficient video understanding in the era of large vision-language models.

Journal

Neural Networks cover
Neural Networks
IF:
6.3
Papers:
7.8K
Citations:
3.0W

Organization

J
Jiangnan University
Scholars:
3.9W
Papers: 2.7W
Citations: 4.7W
U
University of Surrey
Scholars:
1.2W
Papers: 1.3W
Citations: 22