arrow
Return

Efficient and Verified Research Data Extraction with LLM

delete2026-03-13
delete0
PRE
AI
A
Aleksandr Serdiukov
V
Vitaliy Dravgelis
S
Smutin, Daniil *
T
Taldaev, Amir
I
Ivanov, Artem
A
Adonin, Leonid
M
Muravyov, Sergey
DOI:10.3390/a19030214delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large language models (LLMs) hold promise for automated extraction of structured biological information from scientific literature, yet their reliability in some domain-specific tasks, such as DNA probe parsing remains underexplored. We developed a verification-focused, schema-guided extraction pipeline that transforms unstructured texts from scientific articles into a normalized database of oligonucleotide probes, primers, and associated metadata. The system combined multi-turn JSON generation, strict schema validation, sequence-specific rule checks, and a post-processing recovery module that rescues systematically corrupted nucleotide outputs. Benchmarking across nine contemporary LLMs revealed distinct accuracy-hallucination trade-offs, with the context-optimized Qwen3 model achieving the highest overall extraction efficiency while maintaining low hallucination rates. Iterative prompting substantially improved fidelity but introduced notable latency and variance. Across all models, stable error profiles and the success of the recovery module indicated that most extraction failures stem from systematic and correctable formatting issues rather than semantic misunderstandings. These findings highlight both the potential and the current limitations of LLMs for structured scientific data extraction. The research provides a reproducible benchmark and extensible framework for future large-scale curation of molecular biology datasets.
Keywords:
large language model (LLM)
data extraction
nucleotide probe

Journal

Algorithms cover
Algorithms
IF:
2.1
Papers:
631
Citations:
5.4K

Organization

I
itmo university
Scholars:
817
Papers: 283
Citations: 0