arrow
Return

Fully Convolutional CaptionNet: Siamese Difference Captioning Attention Model

delete2019-01-01
delete21
delete
OA
AI
E
Enoch Frimpong
M
Muhammad Umar Aftab
E
Edward Yellakuor Baagyere
Z
Zhiguang Qin *
K
Kifayat Ullah
DOI:10.1109/ACCESS.2019.2957513delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
The generation of the textual description of the differences in images is a relatively new concept that requires the fusion of both computer vision and natural language techniques. In this paper, we present a novel Fully Convolutional CaptionNet (FCC) that employs an encoder-decoder framework to perform visual feature extractions, compute the feature distances, and generate new sentences describing the measured distances. After extracting the features of the images, a contrastive function is used to compute their weighted L1 distance which is learned and selectively attended to determine salient sections of the feature at every time step. The attended feature region is adequately matched to corresponding words iteratively until a sentence is completed. We propose the application of upsampling network to enlarge the features' field of view, this provides a robust pixel-based discrepancy computation. Our extensive experiments indicate that the FCC model outperforms other learning models on the benchmark Spot-the-Diff datasets by generating succinct and meaningful textual differences in images.
Keywords:
Image captioning
deep learning
Siamese network
recurrent neural network
convolutional neural network
attention
fully convolutional networks
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

No organization information available