arrow
返回

Balancing Data on Deep Learning-Based Proteochemometric Activity Classification

delete2021-03-29
delete8
delete
OA
AI
A
Angela Lopez-del Rio *
S
Sergio Picart‐Armada
A
Alexandre Perera-Lluna
DOI:10.1021/acs.jcim.1c00086delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
In silico analysis of biological activity data has become an essential technique in pharmaceutical development. Specifically, the so-called proteochemometric models aim to share information between targets in machine learning ligand-target activity prediction models. However, bioactivity data sets used in proteochemometric modeling are usually imbalanced, which could potentially affect the performance of the models. In this work, we explored the effect of different balancing strategies in deep learning proteochemometric target-compound activity classification models while controlling for the compound series bias through clustering. These strategies were (1) no_resampling, (2) resampling_after_clustering, (3) resampling_before_clustering, and (4) semi_resampling. These schemas were evaluated in kinases, GPCRs, nuclear receptors, and proteases from BindingDB. We observed that the predicted proportion of positives was driven by the actual data balance in the test set. Additionally, it was confirmed that data balance had an impact on the performance estimates of the proteochemometric model. We recommend a combination of data augmentation and clustering in the training set (semi_resampling) to mitigate the data imbalance effect in a realistic scenario. The code of this analysis is publicly available at https://github.com/b2slab/imbalance_pcm_benchmark.
Keyword:
NETWORKS
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Journal of Chemical Information and Modeling 封面图
Journal of Chemical Information and Modeling
IF:
5.3
论文数:
9.1K
被引数:
4.0W

机构

U
universitat politecnica de catalunya
学者数:
1.9W
论文数: 1.6W
被引数: 17