arrow
Return

Korean Erroneous Sentence Classification With Integrated Eojeol Embedding

delete2021-01-01
delete1
delete
OA
AI
D
Donghyun Choi *
I
Ilnam Park
M
Myeongcheol Shin
E
Eunggyun Kim
D
Dong Ryeol Shin
DOI:10.1109/ACCESS.2021.3085864delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
This paper attempts to analyze the Korean sentence classification system. Sentence classification is the task of classifying an input sentence based on predefined categories. However, spelling or space error contained in the input sentence causes problems in morphological analysis and tokenization. This paper proposes a novel approach of Integrated Eojeol (Korean syntactic word separated by space) Embedding to reduce the effect of poorly analyzed morphemes on sentence classification. The paper also proposes two noise insertion methods that further improve classification performance. Our evaluation results indicate that by applying the proposed methods on the existing sentence classifiers, the sentence classification accuracy on erroneous sentences is increased by 8% to 15%.
Keywords:
Training
Task analysis
Tokenization
Licenses
Noise measurement
Network architecture
Encoding
Natural language processing
sentence classification
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

S
sungkyunkwan university (skku)
Scholars:
3.7W
Papers: 3.6W
Citations: 49
K
kakao
Scholars:
74
Papers: 47
Citations: 0