arrow
返回

Language-Guided Multi-Granularity Context Aggregation for Temporal Sentence Grounding

delete2023-01-01
delete4
PRE
AI
G
Guoqiang Gong
L
Linchao Zhu
Y
Yadong Mu *
DOI:10.1109/TMM.2022.3222664delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Temporal sentence grounding in videos is a crucial task in vision-language learning. Its goal is retrieving a video segment from an untrimmed video that semantically corresponds to a natural language query. A video usually contains multiple semantic events, which are rarely isolated. They tend to be temporally ordered and semantically correlated (e.g., some event is often the precursor of another event). To precisely localize a semantic moment from a video, it is critical to effectively extract and aggregate multi-granularity contextual information, including the fine-grained local context around the moment-related video segment (in short snippet-level) and coarse-grained semantic correlation (in segment-level). Additionally, a second main insight in this work is that the above context aggregation should be favorably guided by the queries, rather than fully query-agnostic. Putting above ideas together, we here present a new network that does language-guided multi-granularity context aggregation. It is comprised of two major modules. The core of the first module is a novel language-guided temporal adaptive convolution (LTAC) devised to extract fine-grained information over video snippets around the ground-truth video segment. It decomposes a convolution into two channel-oriented / temporal-oriented ones. In particular, the convolutional channels are supposed to be more susceptible to queries, thus we learn to generate a dynamic channel-oriented kernel with respect to the querying sentence. As a second module, we propose a language-guided global relation block (LGRB) that extracts video-level context. It augments the contextual feature by using a multi-scale temporal attention that tackles the scale variation of ground-truth video segments, and a multi-modal semantic attention that relies on syntactic of the query. For the validation purpose, we have conducted comprehensive experiments on two popularly-adopted video benchmarks (i.e., ActivityNet Captions and Charades-STA). All experimental results and ablation studies have clearly corroborated the effectiveness of our model designs, outstripping prior state-of-the-art methods in terms of major performance metrics for the task.
Keyword:
Videos
Proposals
Grounding
Semantics
Location awareness
Convolution
Task analysis
Vision-language learning
video understanding
temporal sentence grounding
multi-modality learning

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

P
peking university
学者数:
11.9W
论文数: 8.7W
被引数: 146
Z
zhejiang university
学者数:
17.7W
论文数: 12.1W
被引数: 152
引用论文

引用论文

Relationship Between Amyloid β Protein and Melatonin Metabolite in a Study of Electric Utility Workers
err2002-08-01
err0
PREAI
errCurtis W. Noonan; John S. Reif; James B. Burch; Travers Y. Ichinose; Michael G. Yost; Kathy Magnusson
err分享
err收藏
Timekeeping in genetically programmed aging
err1993-03-01
err0
PREAI
errP.E. Kloeden; R. Rössler; O.E. Rössler
err分享
err收藏
Modifications of Surface Integrity during the Cutting of Copper
err2004-12-31
err0
PREAI
errJ. Prohàszka; J. Dobrànszky; J. Nyiró; M. Horvàth; A. G. Mamalis
err分享
err收藏
Adaptation to novel environments during crop diversification
err2020-08-01
err0
PREAI
errGaia Cortinovis; Valerio Di Vittori; Elisa Bellucci; Elena Bitocchi; Roberto Papa
err分享
err收藏
Serial scanning electron microscopy of anti-PKHD1L1 immuno-gold labeled mouse hair cell stereocilia bundles
err2020-06-17
err0
errOAAI
errMaryna V. Ivanchenko; Marcelo Cicconet; Hoor Al Jandal; Xudong Wu; David P. Corey; Artur A. Indzhykulian
err分享
err收藏
err分享
err收藏
Climate policy: Steps to China's carbon peak气候政策: 迈向中国碳峰值的步骤
err2015-06-17
err0
errOAAI
errZhu Liu; Dabo Guan; Scott Moore; Henry Lee; Jun Su; Qiang Zhang
err分享
err收藏
学者 查看更多内容