arrow
Return

Investigating fine-tuning versus zero-shot learning for general large language models when predicting cancer survival from initial oncology consultation documents

delete2026-04-01
delete0
PRE
AI
P
Phaterpekar, T.
M
Mali, Y.
H
Ho, C.
N
Ng, R. T.
B
Bates, A. T.
N
Nunez, J. -J
DOI:10.1016/j.esmorw.2026.100703delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Background: Unstructured oncology consultation notes contain rich clinical information that may support survival prediction. Open-weight large language models (LLMs) can utilize these notes with zero-shot inference or finetuning, but their relative value for this setting remains unclear. The objective of this study is to evaluate open-weight LLMs for predicting 60-month survival from initial oncology consultation notes, comparing (i) zero-shot performance, (ii) performance after fine-tuning, and (iii) smaller natural language processing models trained on the same dataset in prior work. Materials and methods: We used Meta's Llama models to predict patients' 60-month survival using oncology consultation notes from a dataset of 59 800 patients. We tested both zero-shot and fine-tuning approaches. Metrics included balanced accuracy (BA) and weighted F1. Results: Zero-shot performance was limited. Llama-2-13B performed best among the zero-shot configurations (average performance across prompts: BA 0.596, weighted F1 0.644; performance on Prompt 4: BA 0.766, weighted F1 0.802). Fine-tuning improved performance across models: Llama-2-13B achieved BA 0.842, weighted F1 0.846, area under the receiver operating characteristic curve (AUC) 0.905; Llama-2-7B achieved BA 0.840, weighted F1 0.843, AUC 0.911; Llama-3.1-8B achieved BA 0.829, weighted F1 0.829, AUC 0.881. Performance was numerically similar to smaller models trained on the same task and data. Conclusions: For predicting 60-month survival from initial oncology consultation documents, fine-tuning open-weight LLMs meaningfully improves performance compared with zero-shot use, but does not consistently outperform smaller language models. This may suggest that both fine-tuned LLMs and smaller models merit continued investigation, with the most appropriate approach likely to depend on the outcome of interest, clinical context, and practical considerations such as hardware, privacy, and deployment feasibility.
Keywords:
machine learning
LLM
cancer survival
oncology support
prognosis

Journal

E
ESMO Real World Data and Digital Oncology
IF:
0
Papers:
48
Citations:
0

Organization

B
british columbia cancer agency
Scholars:
5.0K
Papers: 3.4K
Citations: 10
U
University of British Columbia
Scholars:
6.9W
Papers: 6.1W
Citations: 8.6W