1
Return

Smiles-based bioactivity prediction through molecular encoder selection and data augmentation

delete2026-07-06
delete0
delete
OA
AI
J
Ju Hyung Lee *
S
Seongik Choi *
U
Utku Özbulak
J
Joris Vankerschaver
W
Wesley De Neve
DOI:10.1186/s13321-026-01251-0delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Quantitative prediction of inhibitor potency can accelerate early-stage drug discovery. Recently, data-driven approaches have gained widespread interest in drug discovery, as evidenced by a growing number of benchmarking challenges and open competitions. In this context, we developed a machine learning-based methodology that can find the most effective way of predicting IC50 values against ASK1 from SMILES, for “Jump AI(.py) 2025: 3rd AI Drug Discovery Competition”, hosted by the Korea Pharmaceutical and Bio-Pharma Manufacturers Association (KPBMA) on the Dacon platform. Applying our methodology achieved the highest overall predictive performance among all participating teams. Beyond this competition setting, we present a compact SMILES-based modeling workflow comprising (i) a pre-trained encoder, (ii) regression models, (iii) data augmentation, and (iv) hyperparameter tuning. We systematically compared molecular representations from sequence- and graph-based models, including ChemBERTa-2 and MolCLR. Across encoder–regressor combinations, ChemBERTa-77 M-MLM embeddings paired with support vector regression (SVR) yielded the strongest predictive performance. Embedding-level mix-up augmentation and SVR hyperparameter tuning further improved predictive performance. Our findings highlight that careful SMILES preprocessing and encoder selection have a critical influence on IC50 values and provide a reproducible benchmark for single-target bioactivity prediction, thus contributing to a more efficient drug discovery process. Scientific Contribution In this study, we propose a machine learning methodology for predicting the IC50 values of ASK1 inhibitors from SMILES representations, with a systematic comparison of molecular encoders and regression models. Our results show that the use of suitable encoder-regressor pairs together with embedding-level mix-up augmentation improves model generalizability without requiring SMILES-level augmentation. This strategy would be particularly useful for settings with imbalanced labels or limited data, and could be applied more broadly to IC50 prediction for other kinase inhibitors.
Keywords:
ASK1
ChemBERTa
IC50 Regression
SMILES

Journal

Journal of Cheminformatics cover
Journal of Cheminformatics
IF:
5.7
Papers:
1.4K
Citations:
1.1W

Organization

C
center for biosystems and biotech data science
Scholars:
9
Papers: 2
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers