arrow
Return

OGER plus plus : hybrid multi-type entity recognition

delete2019-01-21
delete26
delete
OA
AI
L
Lenz Furrer
A
Anna Jancso
N
Nicola Colic
F
Fabio Rinaldi *
DOI:10.1186/s13321-018-0326-3delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Background: We present a text-mining tool for recognizing biomedical entities in scientific literature. OGER++ is a hybrid system for named entity recognition and concept recognition (linking), which combines a dictionary-based annotator with a corpus-based disambiguation component. The annotator uses an efficient look-up strategy combined with a normalization method for matching spelling variants. The disambiguation classifier is implemented as a feed-forward neural network which acts as a postfilter to the previous step. Results: We evaluated the system in terms of processing speed and annotation quality. In the speed benchmarks, the OGER++ web service processes 9.7 abstracts or 0.9 full-text documents per second. On the CRAFT corpus, we achieved 71.4% and 56.7% F1 for named entity recognition and concept recognition, respectively. Conclusions: Combining knowledge-based and data-driven components allows creating a system with competitive performance in biomedical text mining.
Keywords:
Named entity recognition
Concept recognition
Natural language processing
Machine learning
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Cheminformatics cover
Journal of Cheminformatics
IF:
5.7
Papers:
1.5K
Citations:
1.1W

Organization

U
university of zurich
Scholars:
5.0W
Papers: 4.0W
Citations: 65