arrow
返回

Human-Centric Image Captioning

delete2022-06-01
delete14
PRE
AI
P
Pengbo Wang
T
Tianshu Chu
杨洁 (Jie Yang) *
DOI:10.1016/j.patcog.2022.108545delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In this paper, we propose a new topic, Human-Centric Captioning, to mainly describe the human behavior in an image. Human activities and relationships are the primary objectives of visual understanding in daily applications. However, existing image captioning systems cannot differently treat humans and other objects, which limits the ability to understand and describe diverse human activities. As the first explorer of this new task, we build a novel Human-Centric COCO dataset concentrating on humans. Accordingly, we propose a novel Human-Centric Captioning Model (HCCM) that focuses on human-centric feature hierarchization and sentence generation. Specifically, our model first utilizes human body part level knowledge to hierarchize the image features and then applies a novel three-branch captioning model to process these hierarchical features independently to calibrate the descriptions of human actions. Comprehensive experiments demonstrate that our HCCM achieves the state-of-the-art performance with BLEU-4, CIDEr and SPICE scores of 41.5, 127.3, 23.5 respectively. Dataset and code are publicly available at https://github.com/JohnDreamer/HCCM/. (c) 2022 Elsevier Ltd. All rights reserved.
Keyword:
Human-centric
Image captioning
Feature hierarchization

期刊

Pattern Recognition 封面图
Pattern Recognition
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

S
shanghai jiao tong university
学者数:
15.7W
论文数: 11.7W
被引数: 159
引用论文

引用论文

Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
Enhancing the alignment between target words and corresponding frames for video captioning
err2021-03-01
err41
PREAI
errTu, Yunbin; Zhou, Chang; Guo, Junjun; Gao, Shengxiang; Yu, Zhengtao
err分享
err收藏
Automated Peak Picking and Peak Integration in Macromolecular NMR Spectra Using AUTOPSY
err1998-12-01
err0
PREAI
errReto Koradi; Martin Billeter; Max Engeli; Peter Güntert; Kurt Wüthrich
err分享
err收藏
Learning visual relationship and context-aware attention for image captioning学习视觉关系和上下文感知的图像字幕注意
err2020-02-01
err110
PREAI
errWang, Junbo; Wang, Wei; Wang, Liang; Wang, Zhiyong; Feng, David Dagan; Tan, Tieniu
err分享
err收藏
Script identification in natural scene image and video frames using an attention based Convolutional-LSTM network
err2019-01-01
err89
errOAAI
errBhunia, Ankan Kumar; Konwer, Aishik; Bhunia, Ayan Kumar; Bhowmick, Abir; Roy, Partha P.; Pal, Umapada
err分享
err收藏
A survey and analysis on automatic image annotation图像自动标注研究综述与分析
err2018-07-01
err85
PREAI
errCheng, Qimin; Zhang, Qian; Fu, Peng; Tu, Conghuan; Li, Sen
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Dense semantic embedding network for image captioning
err2019-06-01
err40
PREAI
errXiao, Xinyu; Wang, Lingfeng; Ding, Kun; Xiang, Shiming; Pan, Chunhong
err分享
err收藏
没有更多内容