arrow
Return

In-silico predictive mutagenicity model generation using supervised learning approaches

delete2012-05-15
delete18
delete
OA
AI
A
Abhik Seal *
A
Anurag Passi
J
Jaleel, U. C. Abdul
D
David Wild
DOI:10.1186/1758-2946-4-10delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Background: Experimental screening of chemical compounds for biological activity is a time consuming and expensive practice. In silico predictive models permit inexpensive, rapid virtual screening to prioritize selection of compounds for experimental testing. Both experimental and in silico screening can be used to test compounds for desirable or undesirable properties. Prior work on prediction of mutagenicity has primarily involved identification of toxicophores rather than whole-molecule predictive models. In this work, we examined a range of in silico predictive classification models for prediction of mutagenic properties of compounds, including methods such as J48 and SMO which have not previously been widely applied in cheminformatics. Results: The Bursi mutagenicity data set containing 4337 compounds (Set 1) and a Benchmark data set of 6512 compounds (Set 2) were taken as input data set in this work. A third data set (Set 3) was prepared by joining up the previous two sets. Classification algorithms including Naive Bayes, Random Forest, J48 and SMO with 10 fold cross-validation and default parameters were used for model generation on these data sets. Models built using the combined performed better than those developed from the Benchmark data set. Significantly, Random Forest outperformed other classifiers for all the data sets, especially for Set 3 with 89.27% accuracy, 89% precision and ROC of 95.3%. To validate the developed models two external data sets, AID1189 and AID1194, with mutagenicity data were tested showing 62% accuracy with 67% precision and 65% ROC area and 91% accuracy, 91% precision with 96.3% ROC area respectively. A Random Forest model was used on approved drugs from DrugBank and metabolites from the Zinc Database with True Positives rate almost 85% showing the robustness of the model. Conclusion: We have created a new mutagenicity benchmark data set with around 8,000 compounds. Our work shows that highly accurate predictive mutagenicity models can be built using machine learning methods based on chemical descriptors and trained using this set, and these models provide a complement to toxicophores based methods. Further, our work supports other recent literature in showing that Random Forest models generally outperform other comparable machine learning methods for this kind of application.
Keywords:
Molecular descriptors
Machine learning
Mutagenicity
Random forest
Screening
Toxicophores
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Cheminformatics cover
Journal of Cheminformatics
IF:
5.7
Papers:
1.5K
Citations:
1.1W

Organization

I
indiana university system
Scholars:
4.0W
Papers: 3.5W
Citations: 38
I
Indiana University Bloomington
Scholars:
1.9W
Papers: 1.5W
Citations: 2.8W
C
council of scientific & industrial research (csir) - india
Scholars:
4.7W
Papers: 3.9W
Citations: 37
researcher View more organizations
Cited Papers

Cited Papers

Do You See What I Am Saying? Exploring Visual Enhancement of Speech Comprehension in Noisy Environments
err2006-06-13
err0
errOAAI
errL. A. Ross; D. Saint-Amour; V. M. Leavitt; D. C. Javitt; J. J. Foxe
errShare
errSave
Virtual screening of Chinese herbs with random forest
err2007-01-09
err63
PREAI
errEhrman, Thomas M.; Barlow, David J.; Hylands, Peter J.
errShare
errSave
Diffusion across interfaces
err1952-01-01
err0
PREAI
errA. F. H. Ward; L. H. Brooks
errShare
errSave
Thrombin/Prothrombin Interactions with Very Low Density Lipoproteinsa
err2006-12-16
err0
PREAI
errWILLIAM A. BRADLEY; JIAN‐NAN SONG; SANDRA H. GIANTURCO
errShare
errSave
The Effects of Masseter Muscle Paralysis on Facial Bone Growth
err2007-05-01
err0
PREAI
errDamir B. Matic; Arjang Yazdani; R. Glenn Wells; Ting Y. Lee; Bing S. Gan
errShare
errSave
Structure of Polymeric Carbon DioxideCO2−V
err2012-03-19
err0
errOAAI
errFrédéric Datchi; Bidyut Mallick; Ashkan Salamat; Sandra Ninet
errShare
errSave
Bayesian network classifiers
err1997-01-01
err3.8K
errOAAI
errFriedman, N; Geiger, D; Goldszmidt, M
errShare
errSave
researcher View more