arrow
Return

Graph-based Turkish text normalization and its impact on noisy text processing

delete2022-11-01
delete3
delete
OA
AI
Ş
Şeniz Demir *
B
Berkay Topçu
DOI:10.1016/j.jestch.2022.101192delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
User generated texts on the web are freely-available and lucrative sources of data for language technol-ogy researchers. Unfortunately, these texts are often dominated by informal writing styles and the lan-guage used in user generated content poses processing difficulties for natural language tools. Experienced performance drops and processing issues can be addressed either by adapting language tools to user generated content or by normalizing noisy texts before being processed. In this article, we propose a Turkish text normalizer that maps non-standard words to their appropriate standard forms using a graph-based methodology and a context-tailoring approach. Our normalizer benefits from both contex-tual and lexical similarities between normalization pairs as identified by a graph-based subnormalizer and a transformation-based subnormalizer. The performance of our normalizer is demonstrated on a tweet dataset in the most comprehensive intrinsic and extrinsic evaluations reported so far for Turkish. In this article, we present the first graph-based solution to Turkish text normalization with a novel context-tailoring approach, which advances the state-of-the-art results by outperforming other publicly available normalizers. For the first time in the literature, we measure the extent to which the accuracy of a Turkish language processing tool is affected by normalizing noisy texts before being pro-cessed. An analysis of these extrinsic evaluations that focus on more than one Turkish NLP task (i.e., part-of-speech tagger and dependency parser) reveals that Turkish language tools are not robust to noisy texts and a normalizer leads to remarkable performance improvements once used as a preprocessing tool in this morphologically-rich language.(c) 2022 Karabuk University. Publishing services by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Keywords:
Text normalization
Turkish
Graph -based representation
Noisy text
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

E
Engineering Science and Technology-An International Journal-JESTECH
IF:
5.4
Papers:
1.3K
Citations:
6.3K

Organization

M
MEF Universitesi
Scholars:
347
Papers: 360
Citations: 0
T
turkcell turkey
Scholars:
20
Papers: 19
Citations: 0