arrow
Return

Parallel encoder-decoder framework for image captioning

delete2023-12-01
delete3
PRE
AI
P
Peyman Adibi *
H
Hossein Karshenas
A
Alireza Darvishy
DOI:10.1016/j.knosys.2023.111056delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent progress in deep learning has led to successful utilization of encoder-decoder frameworks inspired by machine translation in image captioning models. The stacking of layers in encoders and decoders has made it possible to use several modules in encoders and decoders. However, just one type of module in encoder or decoder has been used in stacked models. In this research, we propose a parallel encoder-decoder framework that aims to take advantage of multiple of types modules in encoders and decoders, simultaneously. This framework contains augmented parallel blocks, which include stacking modules or non-stacked ones. Then, the results of the blocks are integrated to extract higher-level semantic concepts. This general idea is not limited to image captioning and can be customized for many applications that utilize encoder-decoder frameworks. We evaluated our proposed method on the MS-COCO dataset and achieved state-of-the-art results. We got 149.92 for CIDEr-D metric outperforming state-of-the-art image captioning models.
Keywords:
Parallelization
Encoder-decoder framework
Image captioning
Natural language processing

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

U
University of Isfahan
Scholars:
4.5K
Papers: 4.1K
Citations: 5
Z
Zurich University of Applied Sciences
Scholars:
2.2K
Papers: 1.6K
Citations: 2