arrow
Return

A Web Semantic-Based Text Analysis Approach for Enhancing Named Entity Recognition Using PU-Learning and Negative Sampling

delete2023-12-29
delete0
delete
OA
AI
S
Shunqin Zhang
张三国 cover
张三国 (Sanguo Zhang)
W
Wenduo He
张轩 (Xuan Zhang) *
DOI:10.4018/IJSWIS.335113delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The NER task is largely developed based on well-annotated data. However, in many scenarios, the entities may not be fully annotated, leading to serious performance degradation. To address this issue, the authors propose a robust NER approach that combines a novel PU-learning algorithm and negative sampling. Unlike many existing studies, the proposed method adopts a two-step procedure for handling unlabeled entities, thereby enhancing its capability to mitigate the impact of such entities. Moreover, this algorithm demonstrates high versatility and can be integrated into any token-level NER model with ease. The effectiveness of the proposed method is verified on several classic NER models and datasets, demonstrating its strong ability to handle unlabeled entities. Finally, the authors achieve competitive performances on synthetic and real-world datasets.
Keywords:
Negative Sampling
NER
PU-Learning
Robustness
Self-Denoising
Token-Level
Two-Step Procedure
Unlabeled Entity Problem

Journal

I
International Journal on Semantic Web and Information Systems
IF:
5.6
Papers:
471
Citations:
914

Organization

U
university of chinese academy of sciences, cas
Scholars:
4.1W
Papers: 3.8W
Citations: 75
C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704