arrow
返回

NumCap: A Number-controlled Multi-caption Image Captioning Network

delete2023-02-27
delete12
PRE
AI
A
Amr Abdussalam *
Z
Zhongfu Ye
A
Ammar Hawbani
M
Majjed Al-Qatf
R
Rashid Khan
DOI:10.1145/3576927delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Image captioning is a promising task that attracted researchers in the last few years. Existing image captioning models are primarily trained to generate one caption per image. However, an image may contain rich contents, and one caption cannot express its full details. A better solution is to describe an image with multiple captions, with each caption focusing on a specific aspect of the image. In this regard, we introduce a new number-based image captioning model that describes an image with multiple sentences. An image is annotated with multiple ground-truth captions; thus, we assign an external number to each caption to distinguish its order. Given an image-number pair as input, we could achieve different captions for the same image under different numbers. First, a number is attached to the image features to form an image-number vector (INV). Then, this vector and the corresponding caption are embedded using the order-embedding approach. Afterward, the INV's embedding is fed to a language model to generate the caption. To show the efficiency of the numbers incorporation strategy, we conduct extensive experiments using MS-COCO, Flickr30K, and Flickr8K datasets. The proposed model attains 24.1 in METEOR on MS-COCO. The achieved results demonstrate that our method is competitive with a range of state-of-the-art models and validate its ability to produce different descriptions under different given numbers.
Keyword:
Numbers incorporation strategy
encoder-decoder framework
image captioning
order-embedding

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

U
university of science & technology of china, cas
学者数:
3.2W
论文数: 2.7W
被引数: 74
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

Liquid structure of the alkaline-earth metals
err1993-06-01
err0
PREAI
errL. E. González; A. Meyer; M. P. Iñiguez; D. J. González; M. Silbert
err分享
err收藏
Multi-Level Policy and Reward-Based Deep Reinforcement Learning Framework for Image Captioning
err2020-05-01
err78
PREAI
errXu, Ning; Zhang, Hanwang; Liu, An-An; Nie, Weizhi; Su, Yuting; Nie, Jie; Zhang, Yongdong
err分享
err收藏
Accurate online video tagging via probabilistic hybrid modeling
err2014-08-13
err15
PREAI
errShen, Jialie; Wang, Meng; Chua, Tat-Seng
err分享
err收藏
Re-Caption: Saliency-Enhanced Image Captioning Through Two-Phase Learning
err2020-01-01
err52
PREAI
errZhou, Lian; Zhang, Yuejie; Jiang, Yu-Gang; Zhang, Tao; Fan, Weiguo
err分享
err收藏
Image Captioning Model Using Part-of-Speech Guidance Module for Description With Diverse Vocabulary
err2022-01-01
err7
errOAAI
errBae, Ju-Won; Lee, Soo-Hwan; Kim, Won-Yeol; Seong, Ju-Hyeon; Seo, Dong-Hoan
err分享
err收藏
学者 查看更多内容