arrow
返回

EAES: Effective Augmented Embedding Spaces for Text-Based Image Captioning

delete2022-01-01
delete5
delete
OA
AI
K
Khang Nguyen *
D
Doanh C. Bui
T
Truc Trinh
N
Nguyen D. Vo
DOI:10.1109/ACCESS.2022.3158763delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Text-based Image Captioning has been a novel problem since 2020. This topic remains challenging because it requires the model to comprehend not only the visual context but also the scene texts that appear in an image. Therefore, the ways image and scene texts are embedded into the main model for training is crucial. Based on the M4C-Captioner model, this paper proposes the simple but effective EAES embedding module for effectively embedding images and scene texts into the multimodal Transformer layers. In detail, our EAES module contains two significant sub-modules: Objects-augmented and Grid features augmentation. With the Objects-augmented module, we provide the relative geometry feature, representing the relation between objects and between OCR tokens. Furthermore, we extract the grid features for an image with the Grid features augmentation module and combine it with visual objects, which help the model focus on both salient objects and the general context of an image, leading to better performance. We use the TextCaps dataset as the benchmark to prove the effectiveness of our approach on five standard metrics: BLEU4, METEOR, ROUGE-L, SPICE and CIDEr. Without bells and whistles, our method achieves 20.21% on the BLEU4 metric and 85.78% on the CIDEr metric, 1.31% and 4.78% higher, respectively, than the baseline M4C-Captioner method. Furthermore, the results are incredibly competitive with other methods on METEOR, ROUGE-L and SPICE metrics. Source code is available at https://github.com/UIT-Together/EAES_m4c.
Keyword:
Optical character recognition software
Visualization
Feature extraction
Adaptation models
Transformers
Training
Semantics
Image captioning
text-based image captioning
bottom-up top-down
grid feature
multimodal transformer
m4c

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

V
vietnam national university ho chi minh city (vnuhcm) system
学者数:
7.2K
论文数: 4.2K
被引数: 8
引用论文

引用论文

err分享
err收藏
Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation
err2020-01-01
err17
errOAAI
errCheng, Ling; Wei, Wei; Mao, Xianling; Liu, Yong; Miao, Chunyan
err分享
err收藏
Psychosocial impact of illness intrusiveness moderated by self-concept and age in end-stage renal disease.
err1997-01-01
err0
PREAI
errGerald M. Devins; Heather Beanlands; Henry Mandin; Leendert C. Paul
err分享
err收藏
err分享
err收藏
Propiedades psicométricas de una versión breve del Driving Anger Expression Inventory en conductores españoles
err2019-05-28
err0
errOAAI
errDavid Herrero Fernández; Mireia Oliva-Macías; Pamela Parada-Fernández
err分享
err收藏
学者 查看更多内容