arrow
返回

Predicting Molecule Toxicity via Descriptor-Based Graph Self-Supervised Learning

delete2023-01-01
delete3
delete
OA
AI
X
Xinze Li
I
Ilya Makarov *
D
Dmitrii Kiselev *
DOI:10.1109/ACCESS.2023.3308203delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Predicting molecular properties with Graph Neural Networks (GNNs) has recently drawn a lot of attention, with compound toxicity prediction being one of the biggest challenges. In cases where there is insufficient labeled molecule data, an effective approach is to pre-train GNNs on large-scale unlabeled molecular data and then fine-tune them for downstream tasks. Among pre-training strategies, node-level pre-training involves masking and predicting atom properties, while motif-based methods capture rich information in subgraphs. These approaches have shown effectiveness across various downstream tasks. However, current pre-training frameworks face two main challenges: (1) node-level auxiliary tasks do not preserve useful domain knowledge, and (2) the fusion of motif-based methods and node-level tasks is computationally extensive. To address these challenges, we propose Descriptor-based Graph Self-supervised Learning (DGSSL), a method that utilizes domain knowledge to enhance graph representation learning. We extract domain knowledge from a descriptor language known as fragmentary code of substructure superposition (FCSS), where molecules are described using substructures that can serve as centers for weak bonds. Specifically, DGSLL identifies descriptor centers in molecules and encodes motif-like information as special atomic numbers in the pre-training tasks. This enables node-level self-supervised pre-training frameworks for GNNs to also capture rich information in local subgraphs. Experimental results demonstrate that our method achieves state-of-the-art performance on three toxicity-related benchmarks and show their significance in an ablation experiment.
Keyword:
Graphs
molecule graphs
graph neural networks
molecule toxicity prediction
self-supervised learning

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

引用论文

引用论文

Temporal network embedding framework with causal anonymous walks representations
err2022-01-20
err18
errOAAI
errMakarov, Ilya; Savchenko, Andrey; Korovko, Arseny; Sherstyuk, Leonid; Severin, Nikita; Kiselev, Dmitrii; Mikheev, Aleksandr; Babaev, Dmitrii
err分享
err收藏
err分享
err收藏
Interpretation of aeromagnetic data in the Franklin area, northern New Zealand
err2017-01-09
err0
PREAI
errD. J. Robertson; P. N. P. Vidanovich; S. K. Zoellner; J. B. Meyers
err分享
err收藏
Cause of death in patients attending multiple sclerosis clinics
err1991-08-01
err0
PREAI
errA. D. Sadovnick; K. Eisen; G. C. Ebers; D. W. Paty
err分享
err收藏
学者 查看更多内容