返回
STAN: A sequential transformation attention-based network for scene text recognition
DOI:10.1016/j.patcog.2020.107692.png)
摘要
En 中文
Scene text with an irregular layout is difficult to recognize. To this end, a Sequential Transformation Attention-based Network (STAN), which comprises a sequential transformation network and an attention-based recognition network, is proposed for general scene text recognition. The sequential transformation network rectifies irregular text by decomposing the task into a series of patch-wise basic transformations, followed by a grid projection submodule to smooth the junction between neighboring patches. The entire rectification process is able to be trained in an end-to-end weakly supervised manner, requiring only images and their corresponding groundtruth text. Based on the rectified images, an attention-based recognition network is employed to predict a character sequence. Experiments on several benchmarks demonstrate the state-of-the-art performance of STAN on both regular and irregular text. (C) 2020 Elsevier Ltd. All rights reserved.
Keyword:
Scene text recognition
Scene text rectification
Optical character recognition
Deep learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
MORAN: A Multi-Object Rectified Attention Network for scene text recognitionMORAN: 用于场景文本识别的多目标校正注意力网络
PATTERN RECOGNITION
IF7.6
Scene text recognition using a Hough forest implicit shape model and semi-Markov conditional random fields
PATTERN RECOGNITION
IF7.6
A blind deconvolution model for scene text detection and recognition in video
PATTERN RECOGNITION
IF7.6
Curved scene text detection via transverse and longitudinal sequence connection
PATTERN RECOGNITION
IF7.6

