返回
Point-Supervised Video Temporal Grounding
DOI:10.1109/TMM.2022.3205404.png)
摘要
En 中文
Given an untrimmed video and a language query, Video Temporal Grounding (VTG) aims to locate the time interval in the video semantically relevant to the query. Existing fully-supervised VTG methods require accurate annotations of temporal boundary, which is time-consuming and expensive to obtain. On the other hand, weakly-supervised VTG methods where only paired videos and queries are available during training lag far behind the fully-supervised ones. In this paper, we introduce point supervision to narrow the performance gap with affordable annotating cost and propose a novel method dubbed Point-Supervised Video Temporal Grounding (PS-VTG). Specifically, an attention-based grounding network is first employed to obtain a language activation sequence (LAS). Then pseudo segment-level label is generated based on the LAS and the given point supervision to assist the training process. In addition, multi-level distribution calibration and cross-modal contrast are framed to obtain discriminative feature representations and precisely highlight the language-relevant video segments. Experiments on three benchmarks demonstrate that our method trained with point supervision can significantly outperform weakly-supervised approaches and achieve comparable performance with fully-supervised ones.
Keyword:
Cross-modal contrast
multi-level distribution calibration
point supervision
video temporal grounding
期刊
IF:
9.7
论文数:
4.5K
被引数:
2.4W
机构
引用论文
Epigenetic Reprogramming During Somatic Cell Nuclear Transfer: Recent Progress and Future Directions

