arrow
Return

COME: Clip-OCR and Master ObjEct for text image captioning

delete2023-08-01
delete7
PRE
AI
G
Gang Lv *
Y
Yining Sun
F
Fudong Nian
M
Maofei Zhu
W
Wenliang Tang
Z
Zhenzhen Hu
DOI:10.1016/j.imavis.2023.104751delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Text image captioning aims to understand the scene text in images for generating image captions. The key chal-lenge of this task is to accurately and comprehensively understand the OCR tokens of scene text. Due to the dual modal of visual and textual features of scene text, expressing the multimodal semantic features of OCR tokens ac-curately is a challenging task. Additionally, since scene text cannot exist independently of specific objects and is always associated with its surroundings, establishing a scene graph centered around OCR tokens is also an impor-tant approach to understand its relationship with other objects in the image. In this paper, we propose a novel model named Clip -OCR and Master ObjEct (dubbed as COME) for text image captioning. First, we introduce a CLIP-OCR module to enhance the multimodal representation of OCR tokens. We separate the OCR representation into visual and textual items and narrow the similarity by contrastive learning. With the assistance of the CLIP -OCR module, we realize correlation alignment between different modes. Next, we propose the concept of master object for each OCR text and purify the OCR-oriented scene graph with it. The master object is defined as the ob-ject to which the OCR is attached, which bridges the semantic relationship between the OCR tokens and the image. We consider the master object as a proxy that connects OCR tokens and other regions in the image. By ex-ploring the master object for each OCR token, we build a purified scene graph based on the master object and then enrich the visual embedding by the Graph Convolution Network (GCN). Furthermore, we cluster the OCR tokens and append the hierarchical information on the input embedding to provide a complete representation. Experiments on the TextCaps validation set and test set demonstrate the effectiveness of the proposed framework.& COPY; 2023 Elsevier B.V. All rights reserved.
Keywords:
Image captioning
Graph convolution network
OCR
LSTM

Journal

Image and Vision Computing cover
Image and Vision Computing
IF:
4.2
Papers:
4.0K
Citations:
6.7K

Organization

H
hefei institutes of physical science, cas
Scholars:
4.5K
Papers: 3.5K
Citations: 4
C
chinese academy of sciences
Scholars:
56.3W
Papers: 44.8W
Citations: 704