1
Return

MK-ensemble: fragment-based multi-kernel ensemble for interpretable structure-activity relationship modeling of steroidal saponins

delete2026-08-05
delete0
delete
OA
AI
G
Guohao Lv
H
Huichao Liu
X
Xiaolei Zhu
S
Shuai Yang
Q
Qingyong Wang *
L
Lichuan Gu *
DOI:10.1186/s13321-026-01270-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Structure-activity relationship (SAR) modeling of natural products presents a persistent methodological challenge: datasets typically contain fewer than 100 compounds, which restricts the use of data-hungry deep learning models, while conventional QSAR approaches lack mechanistic interpretability and function as black boxes. We present MK-Ensemble, a systematic four-stage optimization framework for interpretable, fragment-based SAR modeling with small-sample natural product datasets. The framework integrates multi-kernel support vector regression, hybrid molecular representations, adversarial domain adaptation, and stacking ensemble learning with strictly nested cross-validation. As a validation case study, we applied the framework to a curated dataset of 91 antioxidant compounds–comprising 24 steroidal saponins from Polygonatum cyrtonema and 67 structurally diverse reference compounds–with 128 activity records across DPPH, ABTS, and FRAP assays. We additionally performed applicability domain characterization via Williams plots and descriptor-space distance analysis, Y-randomization testing (500 permutations), and rigorous statistical model comparison using corrected resampled t-tests and Bayesian correlated t-tests. The Stacking Ensemble achieved $$R^2 = 0.846$$ (95% CI 0.78–0.91; RMSE = 0.154; $$Q^2_{\textrm{CV}} = 0.831$$) for DPPH and $$R^2 = 0.920$$ (95% CI 0.87–0.97; RMSE = 0.089; $$Q^2_{\textrm{CV}} = 0.907$$) for ABTS, with corrected resampled t-tests confirming significance over Random Forest baselines (DPPH: $$p = 0.038$$; ABTS: $$p = 0.003$$). Ablation experiments confirmed that each optimization stage–fragment feature integration (+ 20%), domain adaptation (+ 17%), and ensemble stacking (+ 32%)–contributes measurably to final performance. Y-randomization testing (500 permutations) demonstrated that observed performance is extremely unlikely under null label distributions ($$p < 0.002$$). Applicability domain analysis via Williams plots confirmed that the majority of test-set predictions fall within the reliable prediction domain. Integrated Gradients (IG) analysis correctly recovered the well-established SAR pattern that aglycone cores contribute substantially more to predicted antioxidant activity than glycosylated fragments (scores 1.63–1.72 vs. 0.41–0.54, Wilcoxon rank-sum $$p < 0.001$$). Attribution stability was validated across 100 bootstrap replicates (Spearman $$\rho = 0.87 \pm 0.06$$), and convergent evidence was obtained from SHAP analysis (Spearman $$\rho = 0.91$$) and permutation importance testing. MK-Ensemble provides an interpretable SAR modeling framework for small-sample natural product datasets. The framework demonstrates that predictive performance and fragment-level interpretability can be obtained concurrently under data scarcity. Computational network pharmacology and molecular dynamics simulations provide complementary multi-scale support for the fragment-level interpretations, though experimental validation remains necessary for definitive mechanistic conclusions. Current cheminformatics approaches typically treat kernel-based prediction and fragment-based explanation as separate problems; MK-Ensemble unifies multi-kernel learning, fragment attention, adversarial domain adaptation, and stacking ensemble learning into a single interpretable pipeline optimized for small-sample natural product datasets. The framework shows that strong predictive performance and fragment-level mechanistic interpretability can be obtained concurrently when fewer than 100 compounds are available, a regime where deep learning models often struggle. Applied to steroidal saponins, the model recovers the established structure-activity relationship that aglycone cores dominate antioxidant activity over glycosylated fragments, and this computational finding is complemented by network pharmacology and molecular dynamics simulations, which provide multi-scale computational support for the fragment-level interpretations.
Keywords:
Interpretable machine learning
Fragment-based modeling
Multi-kernel ensemble
Stacking ensemble
Steroidal saponins
Structure-activity relationship
Natural products
Applicability domain
Network pharmacology
Molecular dynamics

Journal

Journal of Cheminformatics cover
Journal of Cheminformatics
IF:
5.7
Papers:
1.4K
Citations:
1.1W

Organization

S
School of Artificial Intelligence
Scholars:
628
Papers: 288
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers