arrow
Return

Multi-level semantic-assisted prototype learning for Few-Shot Action Recognition

delete2025-07-01
delete0
PRE
AI
D
Dan Liu *
Q
Qing Xia
F
Fanrong Meng
M
Mao Ye
张建伟 (Jianwei Zhang)
DOI:10.1016/j.neucom.2025.130022delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The Few-Shot Action Recognition (FSAR) task involves recognizing new categories with limited labeled data. The conventional fine-tuning-based adaptation approach is often prone to overfitting and lacks temporal modeling for video data. Moreover, the discrepancy in distribution between meta-training and meta-test sets can also lead to suboptimal performance in few-shot scenarios. This paper introduces a simple yet effective multi-level semantic-assisted prototype learning framework to tackle these challenges. Initially, we leverage CLIP to achieve multimodal adaptation learning and present a multi-level semantic-assisted learning module to enhance the prototypes of different action classes based on semantic information. Additionally, we integrate the lightweight adapters into the CLIP visual encoder to support parameter-efficient transfer learning and improve temporal modeling in videos. Especially, a bias compensation block is employed for feature rectification to mitigate the distribution bias in FSAR stemming from data scarcity. Extensive experiments conducted on five standard benchmark datasets demonstrate the effectiveness of the proposed method.
Keywords:
Few-shot action recognition
Multimodal learning
Semantic-assisted learning
Bias compensation

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

U
Univ Hamburg
Scholars:
535
Papers: 301
Citations: 137
U
Univ Shanghai Sci and Technol
Scholars:
2.1K
Papers: 922
Citations: 331
U
University of Electronic Science and Technology of China
Scholars:
5.5K
Papers: 2.2K
Citations: 4.0W
researcher View more organizations