返回
Standardizing chemical compounds with language models
DOI:10.1088/2632-2153/ace878.png)
摘要
En 中文
With the growing amount of chemical data stored digitally, it has become crucial to represent chemical compounds accurately and consistently. Harmonized representations facilitate the extraction of insightful information from datasets, and are advantageous for machine learning applications. To achieve consistent representations throughout datasets, one relies on molecule standardization, which is typically accomplished using rule-based algorithms that modify descriptions of functional groups. Here, we present the first deep-learning model for molecular standardization. We enable custom standardization schemes based solely on data, which, as additional benefit, support standardization options that are difficult to encode into rules. Our model achieves over 98%
Keyword:
chemoinformatics
molecule standardization
molecular transformer
compound representation
natural language processing
chemical datasets
期刊
M
IF:
4.6
论文数:
1.1K
被引数:
3.4K
机构
暂无机构信息
引用论文
SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules微笑,一种化学语言和信息系统。1.介绍方法和编码规则
Extraction of organic chemistry grammar from unsupervised learning of chemical reactions从化学反应的无监督学习中提取有机化学语法
SCIENCE ADVANCES
IF12.5
Generating Focused Molecule Libraries for Drug Discovery with Recurrent Neural Networks使用递归神经网络生成用于药物发现的聚焦分子库
ACS CENTRAL SCIENCE
IF10.4
A graph-convolutional neural network model for the prediction of chemical reactivity用于预测化学反应性的图卷积神经网络模型
CHEMICAL SCIENCE
IF7.4
Predicting retrosynthetic pathways using transformer-based models and a hyper-graph exploration strategy
CHEMICAL SCIENCE
IF7.4

