arrow
Return

Learning joint relationship attention network for image captioning

delete2023-01-01
delete16
PRE
AI
X
Xiaodong Gu
DOI:10.1016/j.eswa.2022.118474delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Image captioning aims at automatically describing the main content of an image with a complete and natural sentence. Existing attention-based methods often focus on visual features individually, while ignoring relationship information between image features that provides important guidance for generating captions. To alleviate this issue, in this work we propose the new Joint Relationship Attention Network (JRAN) that novelly explores the relationships between the features in the image. Technically, JRAN capitalizes on semantic features as s supplementary to the region features, fully learn two types of relationships, the visual relationships between region features and the visual-semantic relationships between region features and semantic features. Then, JRAN further make a dynamic trade-off between them during outputting the relationship representation. Moreover, we devise a new feature fusion based attention, which can adaptively fuse the region features and previously obtained relationship representation when generating different words. Extensive experiments on MSCOCO and Flickr30k, Flickr8k datasets show the superiority of our JRAN method qualitatively and quantitatively compared with several related state-of-the-art methods. More remarkably, JRAN achieves 28.3% and 58.2% on Flickr30k, and 22.7% and 55.3% on Flickr8k for the metrics BLEU4 and CIDEr.
Keywords:
Visual relationship
Feature relationship network
Joint relationship learning
Image captioning

Journal

Expert Systems with Applications cover
Expert Systems with Applications
IF:
7.5
Papers:
2.9W
Citations:
10.2W

Organization

F
fudan university
Scholars:
11.7W
Papers: 7.7W
Citations: 121