Return
EGFR mutation prediction in NSCLC from CT images via self-supervised disentangled feature learning and multimodal fusion
J
H
Y
W
H
M
J
K
Z
X
H
Z
DOI:10.1186/s12967-026-08778-8.png)
Abstract
En 中文
Computed tomography (CT)-based noninvasive prediction of epidermal growth factor receptor (EGFR) mutation status in non-small cell lung cancer (NSCLC) is highly important for advancing precision diagnosis and treatment. Current detection methods that rely on invasive tissue biopsies present several limitations. This study aimed to develop a model that integrated deep learning and multimodal information for the noninvasive prediction of EGFR-sensitizing mutations (exon 19 deletion and L858R point mutation) in NSCLC patients using contrast-enhanced CT (CE-CT) images. We retrospectively included 466 pathologically confirmed NSCLC patients from two centers (413 from our own center for model development and internal validation, and 53 from another center for independent external validation), and proposed a novel paradigm combining disentangled self-supervised learning (D-SSL) with multimodal gated fusion. The D-SSL encoder was pretrained on CT images and tumor ROI masks from the entire 413-patient internal cohort without using EGFR mutation labels and was subsequently fixed for downstream classification. An adaptive gated fusion network was subsequently constructed to dynamically integrate these deep features with traditional radiomics features and clinical information to predict the EGFR-sensitizing mutation status. In 20 repeated random splits, the D-SSL fusion models achieved mean AUCs of 0.800 (±0.052) and 0.794 (±0.048) on the training set and internal test set, respectively. One-way ANOVA and subsequent Tukey HSD post-hoc test revealed that the performance on the internal test set of the D-SSL fusion models was significantly superior to that of the clinical models, radiomics models, and radiomics-clinical fusion models (all p < 0.001). On the fully independent external validation set, which was not involved in D-SSL pretraining or downstream model development, the D-SSL fusion models maintained robust performance with a mean AUC of 0.680 (±0.038). Ablation studies showed that models using only the shape head or only the texture head achieved mean AUCs of 0.680 (±0.023) and 0.684 (±0.028), respectively, both significantly lower than that of the complete D-SSL fusion models (both p < 0.001), confirming the complementarity of shape and texture features. Subgroup analysis demonstrated that the model achieved accuracies significantly above the random chance level (50%) in female 83.65% (±3.97%), male 68.90% (±3.72%), smoker 85.07% (±4.41%), and nonsmoker 69.21% (±3.44%) subgroups, suggesting that it did not merely rely on clinical features but captured deep imaging phenotypes associated with EGFR-sensitizing mutations. Feature visualization analysis demonstrated that the D-SSL feature space could effectively distinguish among various imaging manifestation types and pathological categories, confirming the model’s ability to capture highly discriminative and clinically meaningful imaging phenotypes. This study developed a framework integrating D-SSL and multimodal fusion for the noninvasive prediction of EGFR-sensitizing mutations in NSCLC. The framework demonstrated strong discrimination in the internal transductive evaluation and retained predictive value in a fully independent external cohort. These findings support the potential of self-supervised representation learning and gated multimodal fusion, while larger multicenter studies using strictly inductive split-before-pretraining designs remain necessary.
Keywords:
Non-small cell lung cancer
Epidermal growth factor receptor
Disentangled self-supervised learning
Gated fusion
Contrast-enhanced computed tomography
Journal
IF:
7.5
Papers:
9.3K
Citations:
3.2W
