Return
Transformers for molecular property prediction: domain adaptation efficiently improves performance
A
M
S
O
D
A
DOI:10.1186/s13321-026-01252-z.png)
Abstract
En 中文
Over the past six years, molecular transformer models have become a part of the computational toolbox for drug discovery. Most existing models are pre-trained on millions to billions of molecules from large-scale unlabeled datasets such as ZINC or ChEMBL. However, the extent to which such large-scale pre-training improves molecular property prediction remains unclear. This study investigates the potential of transformer models for molecular property prediction while addressing their current limitations. We explore strategies to enhance performance, including the influence of pre-training dataset size and the benefits of domain adaptation through chemically informed objectives. Our results show that increasing the pre-training dataset beyond approximately 400–800K molecules does not improve performance across seven datasets covering five ADME endpoints: lipophilicity, permeability, solubility (two datasets), microsomal stability (two datasets), and plasma protein binding. In contrast, applying domain adaptation on a small number of domain-relevant molecules ($$\le 4K$$) using multi-task regression of physicochemical properties significantly improves model performance across all datasets (P-value < 0.001). Furthermore, we find that a model pre-trained on $$\sim$$400K molecules and adapted on a small domain-specific dataset outperforms larger-scale transformer models like MolFormer and performs comparably to MolBERT. Benchmarking these models alongside baseline representations using RDKit descriptors and Morgan fingerprints reveals that incorporating chemically, physically, and topologically informed features consistently leads to superior performance, regardless of whether used with traditional or transformer-based architectures. While traditional practices such as a random forest model with RDKit descriptors remain strong baselines, this study identifies concrete practices that significantly enhance the performance of transformer models. In particular, aligning pre-training and adaptation with chemically meaningful tasks and domain-relevant data offers a promising path forward for future advancements in molecular property prediction. Our models are available on HuggingFace to allow for easy use and adaptation at https://huggingface.co/collections/UdS-LSV/domain-adaptation-molecular-transformers-6821e7189ada6b7d0a5b62d4. Scientific contribution We introduce the first systematic evaluation of domain adaptation strategies for molecular transformer models in predicting key ADME properties, such as lipophilicity, solubility, clearance, etc. Our results demonstrate that combining pre-training with domain adaptation significantly improves predictive performance and generalization across diverse chemical datasets (P-value $$< 1e^{-9}$$). This contribution advances cheminformatics by offering practical guidelines and open resources for developing more accurate and robust molecular transformer models for property prediction.
Keywords:
Molecular transformers
Domain adaptation
Chemistry-aware modeling
Rigorous analysis
Journal
IF:
5.7
Papers:
1.4K
Citations:
1.1W
