arrow
返回

Exploring coherence from heterogeneous representations for OCR image captioning

delete2024-09-06
delete0
PRE
AI
Y
Yao Zhang
Z
Zijie Song *
Z
Zhenzhen Hu
DOI:10.1007/s00530-024-01470-1delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Text-based image captioning is an important task, aiming to generate descriptions based on reading and reasoning the scene texts in images. Text-based image contains both textual and visual information, which is difficult to be described comprehensively. Recent works fail to adequately model the relationship between features of different modalities and fine-grained alignment. Due to the multimodal characteristics of scene texts, the representations of text usually come from multiple encoders of visual and textual, leading to heterogeneous features. Though lots of works have paid attention to fuse features from different sources, they ignore the direct correlation between heterogeneous features, and the coherence in scene text has not been fully exploited. In this paper, we propose Heterogeneous Attention Module (HAM) to enhance the cross-modal representations of OCR tokens and devote it to text-based image captioning. The HAM is designed to capture the coherence between different modalities of OCR tokens and provide context-aware scene text representations to generate accurate image captions. To the best of our knowledge, we are the first to apply the heterogeneous attention mechanism to explore the coherence in OCR tokens for text-based image captioning. By calculating the heterogeneous similarity, we interactively enhance the alignment between visual and textual information in OCR. We conduct the experiments on the TextCaps dataset. Under the same setting, the results show that our model achieves competitive performances compared with the advanced methods and ablation study demonstrates that our framework enhances the original model in all metrics.
Keyword:
Text-based image captioning
Heterogeneous attention
Scene text
TextCaps

期刊

Multimedia Systems 封面图
Multimedia Systems
IF:
3.1
论文数:
2.8K
被引数:
2.7K

机构

H
hefei university of technology
学者数:
2.5W
论文数: 1.7W
被引数: 35
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
Locating RNAs In Situ with FISH-STIC Probes
err2014-09-04
err0
PREAI
errJohn R. Sinnamon; Kevin Czaplinski
err分享
err收藏
The Utility of Combining the IAD and SES Frameworks
err2019-05-07
err0
errOAAI
errDaniel H. Cole; Graham Epstein; Michael D. McGinnis
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
Embedded Heterogeneous Attention Transformer for Cross-Lingual Image Captioning
err2024-01-01
err2
PREAI
errSong, Zijie; Hu, Zhenzhen; Zhou, Yuanen; Zhao, Ye; Hong, Richang; Wang, Meng
err分享
err收藏
Gehirn und Seele
err
IF0
err2021-02-22
err0
PREAI
errPaul Flechsig
err分享
err收藏
学者 查看更多内容