arrow
返回

Feature Extraction Methods in Quantitative StructureActivity Relationship Modeling: A Comparative Study

delete2020-01-01
delete26
delete
OA
AI
S
Shrooq Alsenan *
I
Isra Al-Turaiki
A
Alaaeldin M. Hafez
DOI:10.1109/ACCESS.2020.2990375delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Computational approaches for synthesizing new chemical compounds have resulted in a major explosion of chemical data in the field of drug discovery. The quantitative structure & x2013;activity relationship (QSAR) is a widely used classification and regression method used to represent the relationship between a chemical structure and its activities. This research focuses on the effect of dimensionality-reduction techniques on a high-dimensional QSAR dataset. Because of the multi-dimensional nature of QSAR, dimensionality-reduction techniques have become an integral part of its modeling process. Principal component analysis (PCA) is a feature extraction technique with several applications in exploratory data analysis, visualization and dimensionality reduction. However, linear PCA is inadequate to handle the complex structure of QSAR data. In light of the wide array of current feature-extraction techniques, we perform a comparative empirical study to investigate five feature-extraction techniques: PCA, kernel PCA, deep generalized autoencoder (dGAE), Gaussian random projection (GRP), and sparse random projection (SRP). The experiments are performed on a high-dimensional QSAR dataset, which comprises 6394 features. The transformed low-dimensional dataset is inputted into a deep learning classification model to predict a QSAR biological activity. Three approaches are adopted to validate and measure the proposed techniques: (i) comparing the performance of the classification models, (ii) visualizing the relationship (correlation) between features in the low-dimension Euclidean space, and (iii) validating the proposed techniques using an external dataset. To the best of our knowledge, this study is the first to investigate and compare the aforementioned feature-extraction techniques in QSAR modeling context. The results obtained provide invaluable insights regarding the behavior of different techniques with both negative and positive classes. With linear PCA as a baseline, we prove that the investigated techniques substantially outperform the baseline in multiple accuracy measures and demonstrate useful ways of extracting significant features.
Keyword:
Autoencoder
blood-brain barrier (BBB) permeability
deep generalized autoencoder (dGAE)
dimensioanlity reduction
feature extraction
Gaussian random projection
principal component analysis
quantitative structure-activity relation (QSAR)
sparse random projection
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

K
King Saud University
学者数:
3.4W
论文数: 3.8W
被引数: 815
P
Princess Nourah bint Abdulrahman University
学者数:
8.3K
论文数: 9.6K
被引数: 10
引用论文

引用论文

Simulating and detecting the quantum spin Hall effect in the kagome optical lattice
err2010-11-04
err0
errOAAI
errGuocai Liu; Shi-Liang Zhu; Shaojian Jiang; Fadi Sun; W. M. Liu
err分享
err收藏
err分享
err收藏
Nonlinear principal components analysis: Introduction and application
err2007-01-01
err468
PREAI
errLinting, Marielle; Meulman, Jacqueline J.; Groenen, Patrick J. F.; van der Kooij, Anita J.
err分享
err收藏
学者 查看更多内容