arrow
Return

Batch-transformer for scene text image super-resolution

delete2024-08-29
delete2
delete
OA
AI
Y
Yaqi Sun
X
Xiaolan Xie *
Z
Zhi Li
K
Kai Yang
DOI:10.1007/s00371-024-03598-7delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recognizing low-resolution text images is challenging as they often lose their detailed information, leading to poor recognition accuracy. Moreover, the traditional methods, based on deep convolutional neural networks (CNNs), are not effective enough for some low-resolution text images with dense characters. In this paper, a novel CNN-based batch-transformer network for scene text image super-resolution (BT-STISR) method is proposed to address this problem. In order to obtain the text information for text reconstruction, a pre-trained text prior module is employed to extract text information. Then a novel two pipeline batch-transformer-based module is proposed, leveraging self-attention and global attention mechanisms to exert the guidance of text prior to the text reconstruction process. Experimental study on a benchmark dataset TextZoom shows that the proposed method BT-STISR achieves the best state-of-the-art performance in terms of structural similarity (SSIM) and peak signal-to-noise ratio (PSNR) metrics compared to some latest methods.
Keywords:
Computer vision
Super-resolution
Scene text image
Batch-transformer
Loss function

Journal

Visual Computer cover
Visual Computer
IF:
2.9
Papers:
4.6K
Citations:
6.5K

Organization

H
Hengyang Normal University
Scholars:
1.2K
Papers: 837
Citations: 886
G
Guangxi Normal University
Scholars:
7.7K
Papers: 4.9K
Citations: 5.1K