arrow
Return

Deep learning for video-text retrieval: a review

delete2023-02-23
delete10
PRE
AI
C
Cunjuan Zhu
Q
Qi Jia
陈伟 (Wei Chen)
Y
Yanming Guo
Y
Yu Liu *
DOI:10.1007/s13735-023-00267-8delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Video-Text Retrieval (VTR) aims to search for the most relevant video related to the semantics in a given sentence, and vice versa. In general, this retrieval task is composed of four successive steps: video and textual feature representation extraction, feature embedding and matching, and objective functions. In the last, a list of samples retrieved from the dataset is ranked based on their matching similarities to the query. In recent years, significant and flourishing progress has been achieved by deep learning techniques, however, VTR is still a challenging task due to the problems like how to learn an efficient spatial-temporal video feature and how to narrow the cross-modal gap. In this survey, we review and summarize over 100 research papers related to VTR, demonstrate state-of-the-art performance on several commonly benchmarked datasets, and discuss potential challenges and directions, with the expectation to provide some insights for researchers in the field of video-text retrieval.
Keywords:
Deep learning
Video-text retrieval
Cross-modal representation
Feature matching
Metric learning

Journal

International Journal of Multimedia Information Retrieval cover
International Journal of Multimedia Information Retrieval
IF:
2.9
Papers:
272
Citations:
866

Organization

D
Dalian University of Technology
Scholars:
5.8W
Papers: 4.3W
Citations: 5.5W
N
national university of defense technology - china
Scholars:
1.8W
Papers: 1.4W
Citations: 9