Return
Multicentre performance and consistency of two deep learning models for malignancy probability estimation of incidental pulmonary nodules
R
N
D
H
M
F
E
C
C
DOI:10.1007/s00330-026-12787-y.png)
Abstract
En 中文
A screening-trained deep learning (DL) model (DL1) for pulmonary nodule malignancy probability estimation on CT previously demonstrated good discrimination on a single-centre dataset of incidental nodules. An updated DL model (DL2) was trained on both screening and clinical data. We aimed to test the performance of both models in a multicentre dataset of incidental nodules. A retrospective, multicentre, case-control dataset of incidental nodules was collected, sampled across size buckets (5–10 mm, 10–15 mm, 15–30 mm), aiming for 10 malignant and 20 benign nodules per bucket and centre, resulting in 270 nodules. Performance was assessed using AUCs and specificity at a fixed sensitivity. AUCs were compared using the DeLong method. Both DL models were compared with the Brock model. Consistent discrimination was investigated by centre-stratified analyses. The multicentre dataset contained 269 nodules (89 malignant) from 231 patients. DL1 and DL2 achieved AUCs of 0.74 and 0.72, respectively, versus 0.63 for Brock (both p < 0.01). Using a 10% threshold for the Brock model (sensitivity 77.5%), specificity was 60% for both DL models compared with 44% for the Brock model. Centre-specific AUCs for DL1, DL2, and Brock were 0.76, 0.75, and 0.74 (centre 1), 0.72, 0.71, and 0.55 (centre 2), and 0.73, 0.71, and 0.59 (centre 3), respectively. The DL models outperformed the Brock model on a multicentre, cancer-enriched size-stratified dataset and performed consistently across centres, though prospective validation in representative cohorts with real-world prevalences is needed. The DL model trained with additional clinical data performed equal to the screening-trained model. Question How consistent is the performance of a screening-trained deep learning model across multicentre incidental pulmonary nodules, and does an additionally clinically trained model improve performance? Findings The two deep learning models showed similar discrimination, outperforming the clinically used Brock model in the pooled cancer-enriched dataset, and consistent performance across multiple centres. Clinical relevance The current study provides evidence of consistent model performance across centres with varying CT characteristics and patient populations within one country, supporting the potential of DL models to classify pulmonary nodules across different clinical settings.
Keywords:
Artificial intelligence
Lung
Neoplasms
Solitary pulmonary nodule
Tomography (X-ray computed)
Journal
IF:
4.7
Papers:
1.6K
Citations:
3.9W
