arrow
Return

Temporal prompt guided visual-text-object alignment for zero-shot video captioning

delete2025-12-12
delete0
PRE
AI
李萍 cover
李萍 (Ping Li) *
王涛 cover
王涛 (Tao Wang)
Z
Zeyu Pan
DOI:10.1016/j.cviu.2025.104601delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• Zero-shot video captioning explores the situation where videos and sentences are unpaired. • Temporal prompts are generated for guiding the video captioning process. • A temporal prompt generation module is developed to capture the actions in video to help yield correct verbs. • A visual-text-object alignment module is designed to align the generated captions and the visual content. • Experiments on three benchmarks demonstrate the superiority of the proposed method.

Journal

Computer Vision and Image Understanding cover
Computer Vision and Image Understanding
IF:
3.5
Papers:
428
Citations:
7.3K

Organization

H
Hangzhou Dianzi University
Scholars:
1.3W
Papers: 9.5K
Citations: 7.5K