arrow
Return

Aligned visual semantic scene graph for image captioning

delete2022-09-01
delete16
PRE
AI
S
Shanshan Zhao
L
Lixiang Li *
H
Haipeng Peng
DOI:10.1016/j.displa.2022.102210delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Image captioning is a multi-modal task to describe an image into natural language. Many state-of-the-art methods generally take the encoder-decoder architecture, encode an image by the convolution neural networks, or by the structured semantic scene graph that contains the object, relationship and the attribute information. The image scene graph constructed by the existing scene graph generation models are generally too noisy. To alleviate the phenomenon, we propose a multi-level cross-modal alignment (MCA) module to align the image scene graph with the sentence scene graph at different level. MCA can distill the redundant information of the image scene graph according to the sentence scene graph, and providing the commonsense knowledge for the decoder. Except for the semantic relationships, we take advantage of the bounding boxes with the visual objects to compute the implicit spatial relationships for the detected objects. With the aligned scene graph features and the implicit spatial relationship information, our decoder fused them via the dynamic mixtured attention to translate these features into descriptions. Extensive experiments on the MSCOCO dataset got the promising result compared with the state-of-the-art methods, which verified the effectiveness of our method.
Keywords:
Deep learning
Image captioning
Scene graph

Journal

Displays cover
Displays
IF:
3.4
Papers:
2.2K
Citations:
3.2K

Organization

B
beijing university of posts & telecommunications
Scholars:
1.4W
Papers: 1.2W
Citations: 9