arrow
Return

Multi-label image classification with recurrently learning semantic dependencies

delete2018-12-15
delete18
PRE
AI
L
Long Chen
R
Ronggui Wang
J
Juan Yang *
L
Lixia Xue
M
Min Hu
DOI:10.1007/s00371-018-01615-0delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recognizing multi-label images is a significant but challenging task toward high-level visual understanding. Remarkable success has been achieved by applying CNN-RNN design-based models to capture the underlying semantic dependencies of labels and predict the label distributions over the global-level features output by CNNs. However, such global-level features often fuse the information of multiple objects, leading to the difficulty in recognizing small object and capturing the label co-relation. To better solve this problem, in this paper, we propose a novel multi-label image classification framework which is an improvement to the CNN-RNN design pattern. By introducing the attention network module in the CNN-RNN architecture, the objects features of the attention map are separated by the channels which are further send to the LSTM network to capture dependencies and predict labels sequentially. A category-wise max-pooling operation is then performed to integrate these labels into the final prediction. Experimental results on PASCAL2007 and MS-COCO datasets demonstrate that our model can effectively exploit the correlation between tags to improve the classification performance as well as better recognize the small targets.
Keywords:
Multi-label
CNN-RNN
Attention
LSTM
Dependencies
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Visual Computer cover
Visual Computer
IF:
2.9
Papers:
4.6K
Citations:
6.5K

Organization

H
hefei university of technology
Scholars:
2.5W
Papers: 1.7W
Citations: 35