1
Return

Transformers for molecular property prediction: domain adaptation efficiently improves performance

delete2026-07-29
delete0
delete
OA
AI
A
Afnan Sultan
M
Max Rausch-Dupont
S
Shahrukh Rafi Khan
O
Olga Kalinina
D
Dietrich Klakow *
A
Andrea Volkamer *
DOI:10.1186/s13321-026-01252-zdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Over the past six years, molecular transformer models have become a part of the computational toolbox for drug discovery. Most existing models are pre-trained on millions to billions of molecules from large-scale unlabeled datasets such as ZINC or ChEMBL. However, the extent to which such large-scale pre-training improves molecular property prediction remains unclear. This study investigates the potential of transformer models for molecular property prediction while addressing their current limitations. We explore strategies to enhance performance, including the influence of pre-training dataset size and the benefits of domain adaptation through chemically informed objectives. Our results show that increasing the pre-training dataset beyond approximately 400–800K molecules does not improve performance across seven datasets covering five ADME endpoints: lipophilicity, permeability, solubility (two datasets), microsomal stability (two datasets), and plasma protein binding. In contrast, applying domain adaptation on a small number of domain-relevant molecules ($$\le 4K$$) using multi-task regression of physicochemical properties significantly improves model performance across all datasets (P-value < 0.001). Furthermore, we find that a model pre-trained on $$\sim$$400K molecules and adapted on a small domain-specific dataset outperforms larger-scale transformer models like MolFormer and performs comparably to MolBERT. Benchmarking these models alongside baseline representations using RDKit descriptors and Morgan fingerprints reveals that incorporating chemically, physically, and topologically informed features consistently leads to superior performance, regardless of whether used with traditional or transformer-based architectures. While traditional practices such as a random forest model with RDKit descriptors remain strong baselines, this study identifies concrete practices that significantly enhance the performance of transformer models. In particular, aligning pre-training and adaptation with chemically meaningful tasks and domain-relevant data offers a promising path forward for future advancements in molecular property prediction. Our models are available on HuggingFace to allow for easy use and adaptation at https://huggingface.co/collections/UdS-LSV/domain-adaptation-molecular-transformers-6821e7189ada6b7d0a5b62d4. Scientific contribution We introduce the first systematic evaluation of domain adaptation strategies for molecular transformer models in predicting key ADME properties, such as lipophilicity, solubility, clearance, etc. Our results demonstrate that combining pre-training with domain adaptation significantly improves predictive performance and generalization across diverse chemical datasets (P-value $$< 1e^{-9}$$). This contribution advances cheminformatics by offering practical guidelines and open resources for developing more accurate and robust molecular transformer models for property prediction.
Keywords:
Molecular transformers
Domain adaptation
Chemistry-aware modeling
Rigorous analysis

Journal

Journal of Cheminformatics cover
Journal of Cheminformatics
IF:
5.7
Papers:
1.4K
Citations:
1.1W

Organization

M
Medical Faculty
Scholars:
844
Papers: 365
Citations: 1
C
Center for Bioinformatics
Scholars:
7
Papers: 4
Citations: 1
S
saarland university
Scholars:
925
Papers: 411
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers