arrow
Return

Local Correspondence Network for Weakly Supervised Temporal Sentence Grounding

delete2021-01-01
delete47
PRE
AI
W
Wenfei Yang
张天柱 (Tianzhu Zhang) *
张勇东 (Yongdong Zhang)
吴枫 (Feng Wu)
DOI:10.1109/TIP.2021.3058614delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Weakly supervised temporal sentence grounding has better scalability and practicability than fully supervised methods in real-world application scenarios. However, most of existing methods cannot model the fine-grained video-text local correspondences well and do not have effective supervision information for correspondence learning, thus yielding unsatisfying performance. To address the above issues, we propose an end-to-end Local Correspondence Network (LCNet) for weakly supervised temporal sentence grounding. The proposed LCNet enjoys several merits. First, we represent video and text features in a hierarchical manner to model the fine-grained video-text correspondences. Second, we design a self-supervised cycle-consistent loss as a learning guidance for video and text matching. To the best of our knowledge, this is the first work to fully explore the fine-grained correspondences between video and text for temporal sentence grounding by using self-supervised learning. Extensive experimental results on two benchmark datasets demonstrate that the proposed LCNet significantly outperforms existing weakly supervised methods.
Keywords:
Grounding
Annotations
Two dimensional displays
Training
Feature extraction
Computational modeling
Task analysis
Weakly supervised
temporal sentence grounding
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704