arrow
Return

Generative External Knowledge for Zero-shot Action Recognition

delete2025-06-01
delete0
PRE
AI
J
Jianing Mao
X
Xinhang Xu
X
Xinran Li
X
Xiang Meng
DOI:10.1016/j.eswa.2025.127420delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Zero-Shot Action Recognition (ZSAR) aims to infer new action classes without any samples of those classes. Traditional methods of acquiring language and visual knowledge limit the relevance and richness of such knowledge, complicating the generalization to the recognition of unseen classes. To address this issue, this paper proposes a novel method based on generative external knowledge for ZSAR. Specifically, high-quality objects are generated to serve as language knowledge of video samples using vision-language large-scale models. Meanwhile, class images are generated from class descriptions using text-to-image large-scale models to serve as visual knowledge of classes. Subsequently, a CLIP pre-trained encoder, combined with a unified multi-stream Transformer framework, learns the enriched cross-modal representations through robust interaction between dual independent encoders. Later, this paper presents a novel predictive model based on Generative Knowledge and Multi Prototype CLIP (GK-MP-CLIP), which leverages the quantity of class images to form a multi-prototype representation for classification support. The proposed method fully applies generative visual and language knowledge to construct the relationship between videos and classes, thereby enhancing ZSAR performance. Experimental results demonstrate significant advantages of the proposed method in enhancing the accuracy of zero-shot recognition for object-interactive actions. This indicates that selecting knowledge sources suited to dataset characteristics can improve ZSAR performance.
Keywords:
Zero-shot action recognition
Generative external knowledge
Cross-modal representation

Journal

Expert Systems with Applications cover
Expert Systems with Applications
IF:
7.5
Papers:
3.0W
Citations:
10.2W

Organization

C
china mobile commun grp henan co ltd
Scholars:
1
Papers: 1
Citations: 0
Cited Papers

Cited Papers

Global Semantic Descriptors for Zero-Shot Action Recognition
err2022-01-01
err0
errOAAI
errEstevam, Valter; Laroca, Rayson; Pedrini, Helio; Menotti, David
errShare
errSave
CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning
err2022-10-01
err301
PREAI
errLuo, Huaishao; Ji, Lei; Zhong, Ming; Chen, Yang; Lei, Wen; Duan, Nan; Li, Tianrui
errShare
errSave
High-Resolution Image Synthesis with Latent Diffusion Models
err2022-06-01
err0
errOAAI
errRobin Rombach; Andreas Blattmann; Dominik Lorenz; Patrick Esser; Bjorn Ommer
errShare
errSave
Fine-tuned CLIP Models are Efficient Video Learners
err2023-06-01
err0
errOAAI
errHanoona Rasheed; Muhammad Uzair Khattak; Muhammad Maaz; Salman Khan; Fahad Shahbaz Khan
errShare
errSave
Relative attributes
err2011-11-01
err0
PREAI
errDevi Parikh; Kristen Grauman
errShare
errSave
A Closer Look at Spatiotemporal Convolutions for Action Recognition
err2018-06-01
err0
errOAAI
errDu Tran; Heng Wang; Lorenzo Torresani; Jamie Ray; Yann LeCun; Manohar Paluri
errShare
errSave
Learning Spatiotemporal Features with 3D Convolutional Networks
err2015-12-01
err0
errOAAI
errDu Tran; Lubomir Bourdev; Rob Fergus; Lorenzo Torresani; Manohar Paluri
errShare
errSave
researcher View more