arrow
返回

Adaptive feature fusion for scene text script identification

delete2024-01-08
delete1
PRE
AI
P
Peng, Fuyou
H
Hui Ma
L
Li Liu *
Y
Yue Lu
C
Ching Y. Suen
DOI:10.1007/s11042-023-17986-zdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Script identification is an essential preliminary step in multilingual OCR systems. This paper focuses primarily on tackling the challenging problem of script identification in scene text images, which are usually characterized by low image quality, diverse text styles, and complex backgrounds. Furthermore, script identification becomes a fine-grained classification problem when some scripts share common characters. To address this issue, we propose a novel end-to-end CNN comprising two streams for extracting distinct types of features, namely, visual features and spatial features. In the visual stream, we introduce an enhanced Squeeze-and-Excitation (SE) channel attention mechanism to emphasize valuable features and suppress irrelevant ones. The enhanced SE is composed of squeeze and excitation steps. The squeeze step employs adaptive average pooling for information aggregation. Two 1x1 convolutional layers are used to derive channel weights in the excitation step. In the spatial stream, we perform efficient analysis of the spatial dependencies within the text lines based on LSTM. Finally, we propose an adaptive fusion approach that combines probability vectors from the two streams. Instead of being fixed, the weight assigned to each probability vector is learned during network training. To validate our proposed method, we conduct extensive tests on four publicly available datasets, viz. MLe2e, RRC-MLT2017, SIW-13, and CVSI-2015. Our proposed method achieves accuracies of 97.66%, 90.24%, 96.66%, and 98.44% on these four datasets, respectively, which compare favorably with state-of-the-art methods. The two streams have demonstrated complementarity. Moreover, ablation experiments have been conducted to verify the effectiveness of each component in the proposed method.
Keyword:
Script identification
Enhanced SE
Adaptive fusion
Two streams

期刊

Multimedia Tools and Applications 封面图
Multimedia Tools and Applications
IF:
3
论文数:
2.0W
被引数:
3.2W

机构

N
Nanchang University
学者数:
3.7W
论文数: 2.1W
被引数: 3.7W
E
east china normal university
学者数:
3.1W
论文数: 2.1W
被引数: 25
C
concordia university - canada
学者数:
8.0K
论文数: 8.9K
被引数: 4
学者 查看更多机构
引用论文

引用论文

Integrating Local CNN and Global CNN for Script Identification in Natural Scene Images
err2019-01-01
err45
errOAAI
errLu, Liqiong; Yi, Yaohua; Huang, Faliang; Wang, Kaili; Wang, Qi
err分享
err收藏
Improving on-line handwritten recognition in interactive machine translation
err2014-03-01
err15
errOAAI
errAlabau, Vicent; Sanchis, Alberto; Casacuberta, Francisco
err分享
err收藏
Improving single-photon sources with Stark tuning
err2007-04-25
err0
PREAI
errMark J. Fernée; Halina Rubinsztein-Dunlop; G. J. Milburn
err分享
err收藏
A cultural approach to wetlands restoration to assess its public acceptance
err2018-11-19
err0
PREAI
errJosep Pueyo‐Ros; Anna Ribas; Rosa M. Fraguell
err分享
err收藏
学者 查看更多内容