arrow
Return

Pseudo-labeling with keyword refining for few-supervised video captioning

delete2025-03-01
delete0
delete
OA
AI
李萍 cover
李萍 (Ping Li) *
王涛 cover
王涛 (Tao Wang)
赵新奎 (Xinkui Zhao)
X
Xianghua Xu
M
Mingli Song
DOI:10.1016/j.patcog.2024.111176delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Video captioning generate a sentence that describes the video content. Existing methods always require a number of captions (e.g., 10 or 20) per video to train the model, which is quite costly. In this work, we explore the possibility of using only one or very few ground-truth sentences, and introduce anew task named few-supervised video captioning. Specifically, we propose a few-supervised video captioning framework that consists of lexically constrained pseudo-labeling module and keyword-refined captioning module. Unlike the random sampling in natural language processing that may cause invalid modifications (i.e., edit words), the former module guides the model to edit words using some actions (e.g., copy, replace, insert, and delete) by a pretrained token-level classifier, and then fine-tunes candidate sentences by a pretrained language model. Meanwhile, the former employs the repetition penalized sampling to encourage the model to yield concise pseudo-labeled sentences with less repetition, and selects the most relevant sentences upon a pretrained videotext model. Moreover, to keep semantic consistency between pseudo-labeled sentences and video content, we develop the transformer-based keyword refiner with the video-keyword gated fusion strategy to emphasize more on relevant words. Extensive experiments on several benchmarks demonstrate the advantages of the proposed approach in both few-supervised and fully-supervised scenarios.
Keywords:
Video captioning
Few supervision
Pseudo-labeling
Keyword refiner
Gated fusion
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

H
Hangzhou Dianzi University
Scholars:
1.3W
Papers: 9.6K
Citations: 7.5K
Z
zhejiang university
Scholars:
17.6W
Papers: 12.1W
Citations: 152