arrow
Return

RECODE-Relational Ecological COrpus for Data Extraction

delete2026-04-17
delete0
PRE
AI
V
Vasco Veiga Branco *
P
Pivovarova, Lidia
K
Kari Lintulaakso
L
Luís Correia
B
Baranovicova, Lenka
L
Lahin, Iiris
D
Dias, Francisco
F
Filipe, David
P
Pedro Cardoso
DOI:10.3897/bdj.14.e177365delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Background Ecology, conservation biology and related disciplines are inherently data-based, with the success of many research projects and initiatives (e.g., protected areas, monitoring plans etc.) being directly dependent on the availability of location and trait information on species and populations. Unfortunately, this data is often either non-existent or available only as unstructured text within publications, especially for megadiverse taxa such as many invertebrate orders. With the emergence of large language models, there have been many attempts to automatically parse such data in machine-readable formats with variable success, either using prompt engineering or training models fit-for-purpose through named entity recognition and relation extraction. Model training has proven more efficient for complex data relations, but it needs labelled corpora, i.e. curated training data containing examples of this information for models to statistically learn from. This is a time-consuming process and, to our knowledge, no standard datasets exist upon which to train new and increasingly better models being released at an increasingly fast pace. New information Here we describe RECODE, a manually annotated corpus of ecological and taxonomic literature, aimed at training and fine-tuning models for automated extraction of occurrence and trait data from unstructured text. All documents presented at this stage have been annotated and validated by experts familiar with the traits of the test taxa (spiders and insects).
Keywords:
biodiversity
functional diversity
insects
large language models
machine learning
named entity recognition
natural language processing
occurrence data
species traits
spiders

Journal

B
Biodiversity Data Journal
IF:
1.1
Papers:
203
Citations:
1.9K

Organization

M
masaryk university
Scholars:
2.6K
Papers: 1.0K
Citations: 0
U
University of Helsinki
Scholars:
4.9K
Papers: 2.0K
Citations: 5.1W
U
Universidade de Lisboa
Scholars:
3.3K
Papers: 1.5K
Citations: 1
researcher View more organizations