arrow
Return

NumCap: A Number-controlled Multi-caption Image Captioning Network

delete2023-02-27
delete12
PRE
AI
A
Amr Abdussalam *
Z
Zhongfu Ye
A
Ammar Hawbani
M
Majjed Al-Qatf
R
Rashid Khan
DOI:10.1145/3576927delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Image captioning is a promising task that attracted researchers in the last few years. Existing image captioning models are primarily trained to generate one caption per image. However, an image may contain rich contents, and one caption cannot express its full details. A better solution is to describe an image with multiple captions, with each caption focusing on a specific aspect of the image. In this regard, we introduce a new number-based image captioning model that describes an image with multiple sentences. An image is annotated with multiple ground-truth captions; thus, we assign an external number to each caption to distinguish its order. Given an image-number pair as input, we could achieve different captions for the same image under different numbers. First, a number is attached to the image features to form an image-number vector (INV). Then, this vector and the corresponding caption are embedded using the order-embedding approach. Afterward, the INV's embedding is fed to a language model to generate the caption. To show the efficiency of the numbers incorporation strategy, we conduct extensive experiments using MS-COCO, Flickr30K, and Flickr8K datasets. The proposed model attains 24.1 in METEOR on MS-COCO. The achieved results demonstrate that our method is competitive with a range of state-of-the-art models and validate its ability to produce different descriptions under different given numbers.
Keywords:
Numbers incorporation strategy
encoder-decoder framework
image captioning
order-embedding

Journal

ACM Transactions on Multimedia Computing Communications and Applications cover
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
Papers:
2.0K
Citations:
5.4K

Organization

U
university of science & technology of china, cas
Scholars:
3.2W
Papers: 2.7W
Citations: 74
C
chinese academy of sciences
Scholars:
56.7W
Papers: 45.0W
Citations: 704
Cited Papers

Cited Papers

Liquid structure of the alkaline-earth metals
err1993-06-01
err0
PREAI
errL. E. González; A. Meyer; M. P. Iñiguez; D. J. González; M. Silbert
errShare
errSave
Multi-Level Policy and Reward-Based Deep Reinforcement Learning Framework for Image Captioning
err2020-05-01
err78
PREAI
errXu, Ning; Zhang, Hanwang; Liu, An-An; Nie, Weizhi; Su, Yuting; Nie, Jie; Zhang, Yongdong
errShare
errSave
Accurate online video tagging via probabilistic hybrid modeling
err2014-08-13
err15
PREAI
errShen, Jialie; Wang, Meng; Chua, Tat-Seng
errShare
errSave
Re-Caption: Saliency-Enhanced Image Captioning Through Two-Phase Learning
err2020-01-01
err52
PREAI
errZhou, Lian; Zhang, Yuejie; Jiang, Yu-Gang; Zhang, Tao; Fan, Weiguo
errShare
errSave
An Ensemble of Generation- and Retrieval-Based Image Captioning With Dual Generator Generative Adversarial Network
err2020-01-01
err33
PREAI
errYang, Min; Liu, Junhao; Shen, Ying; Zhao, Zhou; Chen, Xiaojun; Wu, Qingyao; Li, Chengming
errShare
errSave
Image Captioning Model Using Part-of-Speech Guidance Module for Description With Diverse Vocabulary
err2022-01-01
err7
errOAAI
errBae, Ju-Won; Lee, Soo-Hwan; Kim, Won-Yeol; Seong, Ju-Hyeon; Seo, Dong-Hoan
errShare
errSave
researcher View more