1
Return

Multicentre performance and consistency of two deep learning models for malignancy probability estimation of incidental pulmonary nodules

delete2026-08-08
delete0
delete
OA
AI
R
Renate Dinnessen *
N
Noa Antonissen
D
Dré Peeters
H
Hester A. Gietema
M
Michiel T.H.M. Henkens
F
Firdaus A. A. Mohamed Hoesein
E
Ernst T. Scholten
C
Cornelia Schaefer-Prokop
C
Colin Jacobs
DOI:10.1007/s00330-026-12787-ydelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
A screening-trained deep learning (DL) model (DL1) for pulmonary nodule malignancy probability estimation on CT previously demonstrated good discrimination on a single-centre dataset of incidental nodules. An updated DL model (DL2) was trained on both screening and clinical data. We aimed to test the performance of both models in a multicentre dataset of incidental nodules. A retrospective, multicentre, case-control dataset of incidental nodules was collected, sampled across size buckets (5–10 mm, 10–15 mm, 15–30 mm), aiming for 10 malignant and 20 benign nodules per bucket and centre, resulting in 270 nodules. Performance was assessed using AUCs and specificity at a fixed sensitivity. AUCs were compared using the DeLong method. Both DL models were compared with the Brock model. Consistent discrimination was investigated by centre-stratified analyses. The multicentre dataset contained 269 nodules (89 malignant) from 231 patients. DL1 and DL2 achieved AUCs of 0.74 and 0.72, respectively, versus 0.63 for Brock (both p < 0.01). Using a 10% threshold for the Brock model (sensitivity 77.5%), specificity was 60% for both DL models compared with 44% for the Brock model. Centre-specific AUCs for DL1, DL2, and Brock were 0.76, 0.75, and 0.74 (centre 1), 0.72, 0.71, and 0.55 (centre 2), and 0.73, 0.71, and 0.59 (centre 3), respectively. The DL models outperformed the Brock model on a multicentre, cancer-enriched size-stratified dataset and performed consistently across centres, though prospective validation in representative cohorts with real-world prevalences is needed. The DL model trained with additional clinical data performed equal to the screening-trained model. Question How consistent is the performance of a screening-trained deep learning model across multicentre incidental pulmonary nodules, and does an additionally clinically trained model improve performance? Findings The two deep learning models showed similar discrimination, outperforming the clinically used Brock model in the pooled cancer-enriched dataset, and consistent performance across multiple centres. Clinical relevance The current study provides evidence of consistent model performance across centres with varying CT characteristics and patient populations within one country, supporting the potential of DL models to classify pulmonary nodules across different clinical settings.
Keywords:
Artificial intelligence
Lung
Neoplasms
Solitary pulmonary nodule
Tomography (X-ray computed)

Journal

European Radiology cover
European Radiology
IF:
4.7
Papers:
1.6K
Citations:
3.9W

Organization

D
Department of Medical Imaging
Scholars:
629
Papers: 269
Citations: 0
D
department of pathology
Scholars:
1.2K
Papers: 603
Citations: 0
M
maastricht university medical center
Scholars:
186
Papers: 86
Citations: 0
U
university medical center utrecht
Scholars:
1.7K
Papers: 715
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers