arrow
返回

A Transformer-Based Framework for Scene Text Recognition

delete2022-01-01
delete4
delete
OA
AI
P
Prabu Selvam
K
K. Joseph Abraham Sundar
C
Carlos Andrés Tavera Romero
M
Meshal Alharbi
M
Mehbodniya, Abolfazl
J
Julian Webber
S
Sudhakar Sengan *
DOI:10.1109/ACCESS.2022.3207469delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Scene Text Recognition (STR) has become a popular and long-standing research problem in computer vision communities. Almost all the existing approaches mainly adopt the connectionist temporal classification (CTC) technique. However, these existing approaches are not much effective for irregular STR. In this research article, we introduced a new encoder-decoder framework to identify both regular and irregular natural scene text, which is developed based on the transformer framework. The proposed framework is divided into four main modules: Image Transformation, Visual Feature Extraction (VFE), Encoder and Decoder. Firstly, we employ a Thin Plate Spline (TPS) transformation in the image transformation module to normalize the original input image to reduce the burden of subsequent feature extraction. Secondly, in the VFE module, we use ResNet as the Convolutional Neural Network (CNN) backbone to retrieve text image features maps from the rectified word image. However, the VFE module generates one-dimensional feature maps that are not suitable for locating a multi-oriented text on two-dimensional word images. We proposed 2D Positional Encoding (2DPE) to preserve the sequential information. Thirdly, the feature aggregation and feature transformation are carried out simultaneously in the encoder module. We replace the original scaled dot-product attention model as in the standard transformer framework with an Optimal Adaptive Threshold-based Self-Attention (OATSA) model to filter noisy information effectively and focus on the most contributive text regions. Finally, we introduce a new architectural level bi-directional decoding approach in the decoder module to generate a more accurate character sequence. Eventually, We evaluate the effectiveness and robustness of the proposed framework in both horizontal and arbitrary text recognition through extensive experiments on seven public benchmarks including IIIT5K-Words, SVT, ICDAR 2003, ICDAR 2013, ICDAR 2015, SVT-P and CUTE80 datasets. We also demonstrate that our proposed framework outperforms most of the existing approaches by a substantial margin.
Keyword:
Text recognition
Transformers
Optical character recognition
Decoding
Distortion
Image recognition
Image analysis
Connectionist temporal classification
scene text recognition
self-attention
transformer
optical character recognition
deep learning

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

S
shanmugha arts, science, technology & research academy (sastra)
学者数:
3.1K
论文数: 2.5K
被引数: 0
U
Universidad Santiago de Cali
学者数:
414
论文数: 263
被引数: 245
P
Prince Sattam Bin Abdulaziz University
学者数:
6.9K
论文数: 8.9K
被引数: 9.9K
学者 查看更多机构
引用论文

引用论文

Reading scene text with fully convolutional sequence modeling
err2019-04-01
err51
PREAI
errGao, Yunze; Chen, Yingying; Wang, Jinqiao; Tang, Ming; Lu, Hanqing
err分享
err收藏
err分享
err收藏
err分享
err收藏
Infrared focal plane array incorporating silicon IC process compatible bolometer
err1996-01-01
err0
PREAI
errA. Tanaka; S. Matsumoto; N. Tsukamoto; S. Itoh; K. Chiba; T. Endoh; A. Nakazato; K. Okuyama; Y. Kumazawa; M. Hijikawa; H. Gotoh; T. Tanaka; N. Teranishi
err分享
err收藏
Integration of RF-MEMS resonators on submicrometric commercial CMOS technologies
err2008-11-27
err0
PREAI
errJ L Lopez; J Verd; J Teva; G Murillo; J Giner; F Torres; A Uranga; G Abadal; N Barniol
err分享
err收藏
学者 查看更多内容