Return
Interactive image-to-video transfer learning
DOI:10.1016/j.neunet.2025.108354.png)
Abstract
En 中文
Transfer learning from image to video has become a widely adopted strategy in action recognition. Existing mainstream approaches typically fine-tune the entire network after initialising with pre-trained parameters, which inevitably leads to substantial training overhead. Recent efficiency-oriented methods alleviate this by freezing the backbone during optimisation; however, they tend to overly rely on powerful prior knowledge encoded in image models and often neglect how to effectively perform video reasoning with static backbones. In this paper, we propose an efficient image-to-video transfer learning framework, termed SDST (Static-Dynamic & Spatial-Temporal interaction framework), which explicitly enhances the interaction between static and dynamic cues as well as spatial and temporal domains. Specifically, we introduce the Motion Booster Module that extracts motion descriptors from intermediate features and integrates them via the progressive cross-attention mechanism to fuse static and dynamic representations. Furthermore, we propose a lightweight temporal modelling module, named Channel-Aware Multi-Scale Temporal Modelling, characterised by channel awareness, various temporal scales modelling, as well as feature interactions. By enriching the frozen image backbone through these components, our approach effectively bridges the gap between static visual representations and video-based action recognition tasks. Extensive experiments on Something-Something V1&V2, Diving-48, and Kinetics-400 demonstrate that our method not only surpasses State-of-the-Art efficient transfer learning techniques but also outperforms several fully fine-tuned approaches, all without additional bells and whistles. Moreover, we validate the transferability of the proposed SDST framework to vision-language pre-trained models like CLIP, achieving further gains in performance. These results highlight the potential of our method as a general and scalable solution for efficient video understanding in the era of large vision-language models.
Journal
IF:
6.3
Papers:
7.8K
Citations:
3.0W

