arrow
Return

A comprehensive survey on deep-learning-based visual captioning

delete2023-09-21
delete1
PRE
AI
B
Bowen Xin
徐宁 cover
徐宁 (Ning Xu) *
Y
Yingchen Zhai
T
Tingting Zhang
Z
Zimu Lu
刘静 (Jing Liu)
聂为之 cover
聂为之 (Weizhi Nie)
X
Xuanya Li
刘安安 (An-An Liu)
DOI:10.1007/s00530-023-01175-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Generating a description for an image/video is termed as the visual captioning task. It requires the model to capture the semantic information of visual content and translate them into syntactically and semantically human language. Connecting both research communities of computer vision (CV) and natural language processing (NLP), visual captioning presents the big challenge to bridge the gap between low-level visual features and high-level language information. Thanks to recent advances in deep learning, which are widely applied to the fields of visual and language modeling, the visual captioning methods depending on the deep neural networks has demonstrated state-of-the-art performances. In this paper, we aim to present a comprehensive survey of existing deep learning-based visual captioning methods. Relying on the adopted mechanism and technique to narrow the semantic gap, we divide visual captioning methods into various groups. Representative categories in each group are summarized, and their strengths and limitations are discussed. The quantitative evaluations of state-of-the-art approaches on popular benchmark datasets are also presented and analyzed. Furthermore, we provide the discussions on future research directions.
Keywords:
Visual captioning
Deep learning
Survey

Journal

Multimedia Systems cover
Multimedia Systems
IF:
3.1
Papers:
2.7K
Citations:
2.7K

Organization

T
tianjin university
Scholars:
7.9W
Papers: 5.7W
Citations: 88
H
Heilongjiang University
Scholars:
8.3K
Papers: 5.1K
Citations: 6.8K
B
baidu
Scholars:
577
Papers: 470
Citations: 1
researcher View more organizations