arrow
返回

Unsupervised neural domain adaptation for document image binarization

delete2021-11-01
delete17
delete
OA
AI
F
Francisco J. Castellanos *
A
Antonio‐Javier Gallego
J
Jorge Calvo-Zaragoza
DOI:10.1016/j.patcog.2021.108099delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Binarization is a well-known image processing task, whose objective is to separate the foreground of an image from the background. One of the many tasks for which it is useful is that of preprocessing document images in order to identify relevant information, such as text or symbols. The wide variety of document types, alphabets, and formats makes binarization challenging. There are multiple proposals with which to solve this problem, from classical manually-adjusted methods, to more recent approaches based on machine learning. The latter techniques require a large amount of training data in order to obtain good results; however, labeling a portion of each existing collection of documents is not feasible in practice. This is a common problem in supervised learning, which can be addressed by using the socalled Domain Adaptation (DA) techniques. These techniques take advantage of the knowledge learned in one domain, for which labeled data are available, to apply it to other domains for which there are no labeled data. This paper proposes a method that combines neural networks and DA in order to carry out unsupervised document binarization. However, when both the source and target domains are very similar, this adaptation could be detrimental. Our methodology, therefore, first measures the similarity between domains in an innovative manner in order to determine whether or not it is appropriate to apply the adaptation process. The results reported in the experimentation, when evaluating up to 20 possible combinations among five different domains, show that our proposal successfully deals with the binarization of new document domains without the need for labeled data. (c) 2021 Elsevier Ltd. All rights reserved.
Keyword:
Binarization
Machine learning
Domain adaptation
Adversarial training
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Pattern Recognition 封面图
Pattern Recognition
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

U
universitat d'alacant
学者数:
6.9K
论文数: 7.0K
被引数: 12
引用论文

引用论文

Aligning formative and summative assessments: A collaborative action research challenging teacher conceptions
err2013-06-01
err0
PREAI
errJudith T.M. Gulikers; Harm J.A. Biemans; Renate Wesselink; Marjan van der Wel
err分享
err收藏
Text line detection in handwritten documents
err2008-12-01
err95
PREAI
errLouloudis, G.; Gatos, B.; Pratikakis, I.; Halatsis, C.
err分享
err收藏
A selectional auto-encoder approach for document image binarization
err2019-02-01
err103
errOAAI
errCalvo-Zaragoza, Jorge; Gallego, Antonio-Javier
err分享
err收藏
Plasma dopamine in workers exposed to urban stressor
err2007-08-01
err0
PREAI
errGianfranco Tomei; Assuntina Capozzella; Manuela Ciarrocca; Pina Fiore; Maria Valeria Rosati; Maria Fiaschetti; Teodorico Casale; Vincenza Anzelmo; Francesco Tomei; Carlo Monti
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容