arrow
Return

Generating Chinese named entity data from parallel corpora

delete2014-04-25
delete11
PRE
AI
R
Ruiji Fu
秦兵 (Bing Qin)
T
Ting Liu *
DOI:10.1007/s11704-014-3127-5delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Annotating named entity recognition (NER) training corpora is a costly but necessary process for supervised NER approaches. This paper presents a general framework to generate large-scale NER training data from parallel corpora. In our method, we first employ a high performance NER system on one side of a bilingual corpus. Then, we project the named entity (NE) labels to the other side according to the word level alignments. Finally, we propose several strategies to select high-quality auto-labeled NER training data. We apply our approach to Chinese NER using an English-Chinese parallel corpus. Experimental results show that our approach can collect high-quality labeled data and can help improve Chinese NER.
Keywords:
named entity recognition
Chinese named entity
training data generating
parallel corpora

Journal

Frontiers of Computer Science cover
Frontiers of Computer Science
IF:
4.6
Papers:
1.6K
Citations:
2.8K

Organization

H
harbin institute of technology
Scholars:
8.0W
Papers: 6.6W
Citations: 66