Return
High-quality data selection-driven instruction tuning for biomedical large language models
J
L
X
R
DOI:10.1016/j.jbi.2026.105049.png)
Abstract
En 中文
This study presents a novel data selection framework for enhancing the training efficiency of large language models (LLMs) in biomedical natural language processing (NLP) tasks. We focus on critical tasks sourced from the biomedical dataset, encompassing named entity recognition (NER), relation extraction (RE), event extraction (EE), and text classification (TXTCLASS). These tasks encompass a diverse array of challenges in biomedical NLP and correspond to real-world clinical and research needs. Specifically, our approach introduces the Data Selection (DS) score, a metric designed to quantify the extent to which instructions facilitate response generation by comparing model response losses under conditions with and without instructional context. Notably, we employed the Data Selection (DS) method to filter high-quality data, and further fine-tuned the base model on the selected dataset; the resulting model was named BiomedicalLLM. The main experiment confirm the superiority of our approach, which yields an average F1-score gain of 3.3% and the ablation studies suggest the effectiveness of the overall framework. Analysis reveals that DS dynamically redistributes samples based on task properties: it reduces volume for well-represented tasks while increasing diversity for complex tasks, demonstrating adaptive resource optimization beyond random sampling. This work provides a transformative strategy for optimizing LLM training in biomedical NLP, holding significant potential for applications in clinical practice and biomedical research. The model’s open-source repository is available at https://github.com/dielianhua/BiomedicalLLM .
Keywords:
Data Selection
Biomedical NLP
Large Language Models
Instruction Tuning
Task-specific Optimization
Journal
IF:
4.5
Papers:
3.5K
Citations:
1.9W
