返回
Human-Centric Image Captioning
DOI:10.1016/j.patcog.2022.108545.png)
摘要
En 中文
In this paper, we propose a new topic, Human-Centric Captioning, to mainly describe the human behavior in an image. Human activities and relationships are the primary objectives of visual understanding in daily applications. However, existing image captioning systems cannot differently treat humans and other objects, which limits the ability to understand and describe diverse human activities. As the first explorer of this new task, we build a novel Human-Centric COCO dataset concentrating on humans. Accordingly, we propose a novel Human-Centric Captioning Model (HCCM) that focuses on human-centric feature hierarchization and sentence generation. Specifically, our model first utilizes human body part level knowledge to hierarchize the image features and then applies a novel three-branch captioning model to process these hierarchical features independently to calibrate the descriptions of human actions. Comprehensive experiments demonstrate that our HCCM achieves the state-of-the-art performance with BLEU-4, CIDEr and SPICE scores of 41.5, 127.3, 23.5 respectively. Dataset and code are publicly available at https://github.com/JohnDreamer/HCCM/. (c) 2022 Elsevier Ltd. All rights reserved.
Keyword:
Human-centric
Image captioning
Feature hierarchization
期刊
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
Enhancing the alignment between target words and corresponding frames for video captioning
PATTERN RECOGNITION
IF7.6
Learning visual relationship and context-aware attention for image captioning学习视觉关系和上下文感知的图像字幕注意
PATTERN RECOGNITION
IF7.6
Script identification in natural scene image and video frames using an attention based Convolutional-LSTM network
PATTERN RECOGNITION
IF7.6
没有更多内容

