arrow
Return

Explainable Active Learning Framework for Ligand Binding Affinity Prediction

delete2025-12-23
delete0
delete
OA
AI
S
Satya Pratik Srivastava
R
Rohan Gorantla
S
Sharath Krishna Chundru
C
Claire J.R. Winkelman
A
Antonia S. J. S. Mey
R
Rajeev Kumar Singh
DOI:10.1039/D5DD00436Edelete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Active learning (AL) prioritises which compounds to measure next for protein–ligand affinity when assay or simulation budgets are limited. We present an explainable AL framework built on Gaussian process regression and assess how molecular representations; covariance kernels; and acquisition policies affect enrichment across four drug-relevant targets. Using recall of top active compound; we find that dataset identity which is target’s chemical landscape sets the performance ceiling and method choices modulate outcomes rather than overturn them. Fingerprints with simple Gaussian process kernels provide robust; low-variance enrichment; whereas learned embeddings with non-linear kernels can reach higher peaks but with greater variability. Uncertainty-guided acquisition consistently outperforms random selection; yet no single policy is universally optimal; the best choice follows structure-activity relationship (SAR) complexity. To move beyond black-box selection; we integrate SHapley Additive exPlanations (SHAP) to map high-impact fingerprint bits to chemically interpretable fragments over AL cycles; revealing how model focus sharpens onto SAR-relevant motifs. We release an interactive AL run analysis platform with SHAP traces to support reproducibility and target-specific decision making.
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Digital Discovery cover
Digital Discovery
IF:
5.6
Papers:
979
Citations:
1.7K

Organization

No organization information available