arrow
Return

Text-Conditional Visual-Language Alignment for Video Captioning

delete2025-10-06
delete0
PRE
AI
W
Wenhui Jiang
W
Wenbin Guan
H
Haijun Li
Z
Zhizhen Li
Y
Yuming Fang
彭玉鑫 (Yuxin Peng)
X
Xiaowei Zhao
刘杨 (Yang Liu)
DOI:10.1109/TCSVT.2025.3616201delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Video captioning remains a challenging task due to the diverse video content and the complex relationships between visual and textual elements. Recent efforts predominantly focus on multimodal architecture designs trained with paired video-caption data. Nonetheless, the learning paradigm suffers from the “one-to-many” corresponding problem, since one source video is mapped to multiple caption annotations. The difficulty of video captioning is further exacerbated by the poor-written captions, which mislead the captioner with irrelevant information. Essentially, the problem stems from the inadequate alignment between video and caption. In this work, we propose a Text-Conditional Alignment Transformer, which fully exploits the rich information provided by diverse labeled captions, and avoids the impacts of label ambiguity and noise. To alleviate the challenge of the “one-to-many” correspondence, we introduce Text-conditioned Video Encoding, which diversifies the video representation by emphasizing the spatial-temporal visual areas relevant to the given descriptions while filtering out redundant visual information. The refined video representation is well-aligned to match the corresponding text description, and naturally converts the “one-to-many” mapping to “one-to-one” mapping. To deal with the noisy annotations, we propose Quality-aware Caption Decoding. We first dynamically measure the qualities of different captions corresponding to the same video in a reference-free manner. Then the estimated qualities are further utilized as auxiliary signals, guiding the model to perform quality-aligned learning from noisy captions. We conduct extensive experiments on MSR-VTT, MSVD, VATEX and ActivityNet-Entities datasets, and demonstrate their consistent performance improvements compared to state-of-the-arts.
Keywords:
Video captioning
“one-to-many” correspondence
spatial-temporal alignment
caption quality assessment

Journal

IEEE Transactions on Circuits and Systems for Video Technology cover
IEEE Transactions on Circuits and Systems for Video Technology
IF:
11.1
Papers:
612
Citations:
3.1W

Organization

S
sany heavy industry company ltd.
Scholars:
4
Papers: 3
Citations: 0
J
jiangxi university of finance and economics
Scholars:
74
Papers: 36
Citations: 0
P
peking university
Scholars:
11.7W
Papers: 8.7W
Citations: 146
researcher View more organizations