arrow
Return

Multi-granularity semantic relational mapping for image caption

delete2025-03-01
delete0
PRE
AI
N
Nan Gao *
P
Peng Chen
R
Ronghua Liang
G
Guodao Sun
J
Jijun Tang
DOI:10.1016/j.eswa.2024.125847delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In terms of constructing object-relationship descriptions in images, existing image captioning methods incorporate regional semantic features into visual features to enhance the visual representation. However, they neglect the construction of grid semantic features, resulting in a lack of accurate detailed relationships in the generated results. We propose a M ulti-granularity S emantic R elational M apping(MSRM) framework that dynamically extracts image semantic cue features in place of traditional region labeling in order to get rid of the semantic capability limitation of fixed classification labels and construct grid semantic features. MSRM use the Internal Semantic Mapping mechanism to refine semantic features by filtering out irrelevant features and mapping them onto region and grid features. Simultaneously, the Semantic Mapping mechanism is used to integrate the composite features derived from regions and grids, thereby addressing the problem of describing semantic relationships among objects across different granularities. Experiments on the MSCOCO and Flickr30k datasets show that the proposed MSRM significantly outperforms the state-of-the-art baselines by more than 4% in 7 different metrics including BLEUs, Meteor, Rouge and CIDEr.
Keywords:
Image caption
Multi-granularity
Dynamical semantic cue
Cross-attention

Journal

Expert Systems with Applications cover
Expert Systems with Applications
IF:
7.5
Papers:
2.9W
Citations:
10.2W

Organization

Z
zhejiang university of technology
Scholars:
3.2W
Papers: 2.0W
Citations: 22
C
chinese academy of sciences
Scholars:
56.0W
Papers: 44.8W
Citations: 704