arrow
Return

Human-Centric Image Captioning

delete2022-06-01
delete14
PRE
AI
P
Pengbo Wang
T
Tianshu Chu
杨洁 (Jie Yang) *
DOI:10.1016/j.patcog.2022.108545delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In this paper, we propose a new topic, Human-Centric Captioning, to mainly describe the human behavior in an image. Human activities and relationships are the primary objectives of visual understanding in daily applications. However, existing image captioning systems cannot differently treat humans and other objects, which limits the ability to understand and describe diverse human activities. As the first explorer of this new task, we build a novel Human-Centric COCO dataset concentrating on humans. Accordingly, we propose a novel Human-Centric Captioning Model (HCCM) that focuses on human-centric feature hierarchization and sentence generation. Specifically, our model first utilizes human body part level knowledge to hierarchize the image features and then applies a novel three-branch captioning model to process these hierarchical features independently to calibrate the descriptions of human actions. Comprehensive experiments demonstrate that our HCCM achieves the state-of-the-art performance with BLEU-4, CIDEr and SPICE scores of 41.5, 127.3, 23.5 respectively. Dataset and code are publicly available at https://github.com/JohnDreamer/HCCM/. (c) 2022 Elsevier Ltd. All rights reserved.
Keywords:
Human-centric
Image captioning
Feature hierarchization

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

S
shanghai jiao tong university
Scholars:
15.7W
Papers: 11.7W
Citations: 159
Cited Papers

Cited Papers

Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
errShare
errSave
Enhancing the alignment between target words and corresponding frames for video captioning
err2021-03-01
err41
PREAI
errTu, Yunbin; Zhou, Chang; Guo, Junjun; Gao, Shengxiang; Yu, Zhengtao
errShare
errSave
Automated Peak Picking and Peak Integration in Macromolecular NMR Spectra Using AUTOPSY
err1998-12-01
err0
PREAI
errReto Koradi; Martin Billeter; Max Engeli; Peter Güntert; Kurt Wüthrich
errShare
errSave
Learning visual relationship and context-aware attention for image captioning
err2020-02-01
err110
PREAI
errWang, Junbo; Wang, Wei; Wang, Liang; Wang, Zhiyong; Feng, David Dagan; Tan, Tieniu
errShare
errSave
Script identification in natural scene image and video frames using an attention based Convolutional-LSTM network
err2019-01-01
err89
errOAAI
errBhunia, Ankan Kumar; Konwer, Aishik; Bhunia, Ayan Kumar; Bhowmick, Abir; Roy, Partha P.; Pal, Umapada
errShare
errSave
A survey and analysis on automatic image annotation
err2018-07-01
err85
PREAI
errCheng, Qimin; Zhang, Qian; Fu, Peng; Tu, Conghuan; Li, Sen
errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
Dense semantic embedding network for image captioning
err2019-06-01
err40
PREAI
errXiao, Xinyu; Wang, Lingfeng; Ding, Kun; Xiang, Shiming; Pan, Chunhong
errShare
errSave
no more