arrow
Return

Sample-efficient generative molecular design using memory manipulation

delete2026-03-17
delete0
PRE
AI
J
Jeff Guo *
J
Junwu Chen
A
Anthony Guanxun Chen
P
Philippe Schwaller *
DOI:10.1038/s42256-026-01200-4delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Generative molecular design for drug discovery has recently achieved a wave of experimental validation. Language models operating on string-based representations of molecules are amongst the most successful architectures. The most important factor for downstream success is whether an in silico oracle (computational predictor of a molecule property) is well correlated with the desired end point (such as binding affinity). To this end, current methods use cheaper proxy oracles with a higher throughput before evaluating the most promising subset with high-fidelity oracles. The ability to directly generate molecules with optimal properties as predicted by high-fidelity oracles (computationally expensive simulations with greater predictive accuracy) could greatly enhance generative design and improve hit rates. However, current models are not efficient enough to consider such a prospect, exemplifying the sample efficiency problem. Recently, the Mamba architecture has been proposed as an alternative to transformers, which are widely used in large language models. Existing works have validated Mamba’s performance on tasks spanning natural language completion to biology foundation models. In this work, we introduce a framework called Saturn, which demonstrates the application of the Mamba architecture for generative molecular design. Here we elucidate how experience replay with data augmentation improves the sample efficiency and how Mamba intensifies the effect of this mechanism. Next, we show that Mamba with experience replay outperforms 16 models on multiparameter optimization tasks relevant to drug discovery and possesses sufficient sample efficiency to directly optimize density functional theory simulations as a high-fidelity oracle. Guo et al. train a Mamba-based language model for molecule generation and find that data augmentation and experience replay can enable the efficient generation of property-optimized small molecules.
Keywords:
Cheminformatics
Drug discovery
Engineering
general

Journal

Nature Machine Intelligence cover
Nature Machine Intelligence
IF:
23.9
Papers:
1.3K
Citations:
1.5W

Organization

N
new york university
Scholars:
6.1K
Papers: 2.9K
Citations: 1
E
epfl
Scholars:
713
Papers: 272
Citations: 0