1
Return

Extraction of distant recurrence sites for breast cancer patients from free-text clinical notes using large language models

delete2026-05-23
delete0
PRE
AI
M
Madhu Babu Sikha *
A
Amara Tariq
A
Allison W. Kurian
K
Kevin C. Ward
T
T. H. M. Keegan
D
Daniel L. Rubin
I
Imon Banerjee
DOI:10.1016/j.jbi.2026.105032delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Objectives: Accurate documentation of distant recurrence sites in breast cancer is essential for evaluating treatment effectiveness and outcomes research. However, such information is embedded in unstructured clinical notes, making manual abstraction labor-intensive. Large language models (LLMs) offer a scalable solution for extracting complex information from heterogeneous clinical narratives; however, generic LLMs often lack the specialized clinical reasoning needed for accurate interpretation of oncologic documentation. This study aims to develop an efficient LLM-based framework to automatically extract distant recurrence sites from free-text documentation. Materials & methods: We used clinical notes, pathology and radiology reports from recurrent breast cancer patients at Mayo Clinic (n = 766) for model development and evaluated generalizability on internal hold-out samples (n = 112) and an external Stanford Medicine cohort (n = 110). For cross-disease domain adaptation, we further validated on prostate cancer patients (n = 49). Our proposed framework employs BioLinkBERT, a pretrained language model (PLM) backbone, with weak supervision and an epoch-wise entropy optimization to address limited labeled data and class imbalance across recurrence sites. The fine-tuned model was compared against state-of-the-art models, including Llama2-7B, Llama-3-8B and MedAlpaca, using precision, recall, and F1score. Results: The fine-tuned model outperformed generic and domain-specific LLM baselines, with notable gains in identifying multi-site distant recurrence. In-domain validation showed consistent F1-score improvement (average 0.78), particularly for rare recurrence sites. The model also demonstrated strong performance on the external Stanford cohort and on prostate cancer, achieving F1-score of 0.83 and 0.93, respectively. Conclusion: This study presents an efficient, weakly supervised LLM framework that accurately extracts metastatic recurrence sites, reducing reliance on manual chart review. The results demonstrate that relatively small LLMs, optimized with domain-aware weak supervision, can outperform larger models for complex oncologic information extraction. The model is released as a platform-independent Docker image to support seamless cancer registry integration.
Keywords:
Breast cancer
Sites of distant recurrence
Large language models
Natural language processing
Entropy optimization
Weakly-supervised fine-tuning

Journal

Journal of Biomedical Informatics cover
Journal of Biomedical Informatics
IF:
4.5
Papers:
3.5K
Citations:
1.9W

Organization

 
 emory university
Scholars:
1.9K
Papers: 747
Citations: 0
R
rollins school public health
Scholars:
5.2K
Papers: 4.3K
Citations: 8
M
mayo clinic
Scholars:
8.0W
Papers: 6.5W
Citations: 84
M
mayo clinic phoenix
Scholars:
6.9K
Papers: 5.5K
Citations: 4
S
stanford university
Scholars:
9.2K
Papers: 3.6K
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers