arrow
返回

Mixed-Supervised Scene Text Detection With Expectation-Maximization Algorithm

delete2022-01-01
delete14
PRE
AI
M
Mengbiao Zhao
W
Wei Feng
殷飞 (Fei Yin)
X
Xu-Yao Zhang
刘程琳 (Cheng‐Lin Liu) *
DOI:10.1109/TIP.2022.3197987delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Scene text detection is an important and challenging task in computer vision. For detecting arbitrarily-shaped texts, most existing methods require heavy data labeling efforts to produce polygon-level text region labels for supervised training. In order to reduce the cost in data labeling, we study mixed-supervised arbitrarily-shaped text detection by combining various weak supervision forms (e.g., image-level tags, coarse, loose and tight bounding boxes), which are far easier to annotate. Whereas the existing weakly-supervised learning methods (such as multiple instance learning) do not promote full object coverage, to approximate the performance of fully-supervised detection, we propose an Expectation-Maximization (EM) based mixed-supervised learning framework to train scene text detector using only a small amount of polygon-level annotated data combined with a large amount of weakly annotated data. The polygon-level labels are treated as latent variables and recovered from the weak labels by the EM algorithm. A new contour-based scene text detector is also proposed to facilitate the use of weak labels in our mixed-supervised learning framework. Extensive experiments on six scene text benchmarks show that (1) using only 10% strongly annotated data and 90% weakly annotated data, our method yields comparable performance to that of fully supervised methods, (2) with 100% strongly annotated data, our method achieves state-of-the-art performance on five scene text benchmarks (CTW1500, Total-Text, ICDAR-ArT, MSRA-TD500, and C-SVT), and competitive results on the ICDAR2015 Dataset. We will make our weakly annotated datasets publicly available.
Keyword:
Costs
Annotations
Training
Labeling
Detectors
Data models
Benchmark testing
Mixed-supervised learning
scene text detection
weak supervision forms
expectation-maximization algorithm

期刊

IEEE Transactions on Image Processing 封面图
IEEE Transactions on Image Processing
IF:
13.7
论文数:
1.0W
被引数:
8.4W

机构

C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

Blastococcus capsensis sp. nov., isolated from an archaeological Roman pool and emended description of the genus Blastococcus, B. aggregatus, B. saxobsidens, B. jejuensis and B. endophyticus
err2016-11-01
err0
errOAAI
errKarima Hezbri; Moussa Louati; Imen Nouioui; Maher Gtari; Manfred Rohde; Cathrin Spröer; Peter Schumann; Hans-Peter Klenk; Faten Ghodhbane-Gtari; Maria del Carmen Montero-Calasanz
err分享
err收藏
Functionalization, Modification, and Transformation of Platinum Chini Clusters
err2018-07-06
err0
PREAI
errBeatrice Berti; Cristina Femoni; Maria Carmela Iapalucci; Silvia Ruggieri; Stefano Zacchini
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容