Return
Temporal prompt guided visual-text-object alignment for zero-shot video captioning
DOI:10.1016/j.cviu.2025.104601.png)
Abstract
En 中文
• Zero-shot video captioning explores the situation where videos and sentences are unpaired. • Temporal prompts are generated for guiding the video captioning process. • A temporal prompt generation module is developed to capture the actions in video to help yield correct verbs. • A visual-text-object alignment module is designed to align the generated captions and the visual content. • Experiments on three benchmarks demonstrate the superiority of the proposed method.
Journal
IF:
3.5
Papers:
428
Citations:
7.3K

