Return
MK-ensemble: fragment-based multi-kernel ensemble for interpretable structure-activity relationship modeling of steroidal saponins
G
H
X
S
Q
L
DOI:10.1186/s13321-026-01270-x.png)
Abstract
En 中文
Structure-activity relationship (SAR) modeling of natural products presents a persistent methodological challenge: datasets typically contain fewer than 100 compounds, which restricts the use of data-hungry deep learning models, while conventional QSAR approaches lack mechanistic interpretability and function as black boxes. We present MK-Ensemble, a systematic four-stage optimization framework for interpretable, fragment-based SAR modeling with small-sample natural product datasets. The framework integrates multi-kernel support vector regression, hybrid molecular representations, adversarial domain adaptation, and stacking ensemble learning with strictly nested cross-validation. As a validation case study, we applied the framework to a curated dataset of 91 antioxidant compounds–comprising 24 steroidal saponins from Polygonatum cyrtonema and 67 structurally diverse reference compounds–with 128 activity records across DPPH, ABTS, and FRAP assays. We additionally performed applicability domain characterization via Williams plots and descriptor-space distance analysis, Y-randomization testing (500 permutations), and rigorous statistical model comparison using corrected resampled t-tests and Bayesian correlated t-tests. The Stacking Ensemble achieved $$R^2 = 0.846$$ (95% CI 0.78–0.91; RMSE = 0.154; $$Q^2_{\textrm{CV}} = 0.831$$) for DPPH and $$R^2 = 0.920$$ (95% CI 0.87–0.97; RMSE = 0.089; $$Q^2_{\textrm{CV}} = 0.907$$) for ABTS, with corrected resampled t-tests confirming significance over Random Forest baselines (DPPH: $$p = 0.038$$; ABTS: $$p = 0.003$$). Ablation experiments confirmed that each optimization stage–fragment feature integration (+ 20%), domain adaptation (+ 17%), and ensemble stacking (+ 32%)–contributes measurably to final performance. Y-randomization testing (500 permutations) demonstrated that observed performance is extremely unlikely under null label distributions ($$p < 0.002$$). Applicability domain analysis via Williams plots confirmed that the majority of test-set predictions fall within the reliable prediction domain. Integrated Gradients (IG) analysis correctly recovered the well-established SAR pattern that aglycone cores contribute substantially more to predicted antioxidant activity than glycosylated fragments (scores 1.63–1.72 vs. 0.41–0.54, Wilcoxon rank-sum $$p < 0.001$$). Attribution stability was validated across 100 bootstrap replicates (Spearman $$\rho = 0.87 \pm 0.06$$), and convergent evidence was obtained from SHAP analysis (Spearman $$\rho = 0.91$$) and permutation importance testing. MK-Ensemble provides an interpretable SAR modeling framework for small-sample natural product datasets. The framework demonstrates that predictive performance and fragment-level interpretability can be obtained concurrently under data scarcity. Computational network pharmacology and molecular dynamics simulations provide complementary multi-scale support for the fragment-level interpretations, though experimental validation remains necessary for definitive mechanistic conclusions. Current cheminformatics approaches typically treat kernel-based prediction and fragment-based explanation as separate problems; MK-Ensemble unifies multi-kernel learning, fragment attention, adversarial domain adaptation, and stacking ensemble learning into a single interpretable pipeline optimized for small-sample natural product datasets. The framework shows that strong predictive performance and fragment-level mechanistic interpretability can be obtained concurrently when fewer than 100 compounds are available, a regime where deep learning models often struggle. Applied to steroidal saponins, the model recovers the established structure-activity relationship that aglycone cores dominate antioxidant activity over glycosylated fragments, and this computational finding is complemented by network pharmacology and molecular dynamics simulations, which provide multi-scale computational support for the fragment-level interpretations.
Keywords:
Interpretable machine learning
Fragment-based modeling
Multi-kernel ensemble
Stacking ensemble
Steroidal saponins
Structure-activity relationship
Natural products
Applicability domain
Network pharmacology
Molecular dynamics
Journal
IF:
5.7
Papers:
1.4K
Citations:
1.1W
