arrow
Return

Text mining for the biocuration workflow

delete2012-04-18
delete65
delete
OA
AI
L
Lynette Hirschman *
G
Gully Burns
M
Martin Krallinger
C
Cecilia N. Arighi
K
K. B. Cohen
A
Alfonso Valencia
C
Cathy Wu
A
Andrew Chatr‐aryamontri
K
Karen Dowell
E
Eva Huala
A
Anália Lourenço
R
Robert S Nash
A
A.-L. Veuthey
T
T.A. Wiegers
A
Andrew Winter
DOI:10.1093/database/bas020delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Molecular biology has become heavily dependent on biological knowledge encoded in expert curated biological databases. As the volume of biological literature increases, biocurators need help in keeping up with the literature; (semi-) automated aids for biocuration would seem to be an ideal application for natural language processing and text mining. However, to date, there have been few documented successes for improving biocuration throughput using text mining. Our initial investigations took place for the workshop on 'Text Mining for the BioCuration Workflow' at the third International Biocuration Conference (Berlin, 2009). We interviewed biocurators to obtain workflows from eight biological databases. This initial study revealed high-level commonalities, including (i) selection of documents for curation; (ii) indexing of documents with biologically relevant entities (e. g. genes); and (iii) detailed curation of specific relations (e. g. interactions); however, the detailed workflows also showed many variabilities. Following the workshop, we conducted a survey of biocurators. The survey identified biocurator priorities, including the handling of full text indexed with biological entities and support for the identification and prioritization of documents for curation. It also indicated that two-thirds of the biocuration teams had experimented with text mining and almost half were using text mining at that time. Analysis of our interviews and survey provide a set of requirements for the integration of text mining into the biocuration workflow. These can guide the identification of common needs across curated databases and encourage joint experimentation involving biocurators, text mining developers and the larger biomedical research community.
Keywords:
DATABASE
TAVERNA
BIOLIT
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

D
Database-The Journal of Biological Databases and Curation
IF:
3.6
Papers:
1.7K
Citations:
6.1K

Organization

M
MITRE Corporation
Scholars:
484
Papers: 278
Citations: 1
C
centro nacional de investigaciones oncologicas (cnio)
Scholars:
3.1K
Papers: 1.9K
Citations: 1
C
Carnegie Institution for Science
Scholars:
4.3K
Papers: 4.9K
Citations: 1.0W
University of Colorado System cover
University of Colorado System
Scholars:
6.3W
Papers: 5.5W
Citations: 1.8K
U
universidade do minho
Scholars:
1.1W
Papers: 1.1W
Citations: 10
U
university of southern california
Scholars:
4.6W
Papers: 3.8W
Citations: 51
U
University of Delaware
Scholars:
1.3W
Papers: 1.3W
Citations: 2.0W
S
Stanford University
Scholars:
9.6W
Papers: 8.2W
Citations: 17.0W
J
Jackson Laboratory
Scholars:
3.0K
Papers: 2.0K
Citations: 5.5K
U
university of maine orono
Scholars:
2.5K
Papers: 2.1K
Citations: 3
U
University of Edinburgh
Scholars:
5.2W
Papers: 4.6W
Citations: 71
G
Georgetown University
Scholars:
1.6W
Papers: 1.3W
Citations: 1.5W
researcher View more organizations