返回
Statistical characterization and classification of colon microarray gene expression data using multiple machine learning paradigms
DOI:10.1016/j.cmpb.2019.04.008.png)
摘要
En 中文
Objective: A colon microarray data is a repository of thousands of gene expressions with different strengths for each cancer cell. It is necessary to detect which genes are responsible for cancer growth. This study presents an exhaustive comparative study of different machine learning (ML) systems which serves two major purposes: (a) identification of high risk differential genes using statistical tests and (b) development of a ML strategy for predicting cancer genes. Methods: Four statistical tests namely: Wilcoxon sign rank sum (WCSRS), t test, Kruskal-Wallis (KW), and F-test were adapted for cancerous gene identification using their p-values. The extracted gene set was used to classify cancer patients using ten classifiers namely: linear discriminant analysis (LDA), quadratic discriminant analysis (QDA), naive Bayes (NB), Gaussian process classification (GPC), support vector machine (SVM), artificial neural network (ANN), logistic regression (LR), decision tree (DT), Adaboost (AB), and random forest (RF). Performance was then evaluated using cross-validation protocols and standardized metrics viz. accuracy (ACC) and area under the curve (AUC). Results: The colon cancer dataset consists of 2000 genes from 62 patients (40 cancer vs. 22 control). The overall mean ACC of our ML system using all four statistical tests and all ten classifiers was 90.50%. The ML system showed an ACC of 99.81% using a combination WCSRS test and RF-based classifier. This is an improvement of 8% over previously published values in literature. Conclusions: RF-based model with statistical tests for detection of high risk genes showed the best performance for accurate cancer classification in multi-center clinical trials. (C) 2019 Elsevier B.V. All rights reserved.
Keyword:
Colon cancer
Gene expression data
Prediction
Statistical test
Machine learning
Performance
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
4.8
论文数:
7.0K
被引数:
2.1W
机构
引用论文
Tumor tissue identification based on gene expression data using DWT feature extraction and PNN classifier
NEUROCOMPUTING
IF6.5
Stability of Iodine in Iodized Salt Used for Correction of Iodine-Deficiency Disorders. II用于纠正碘缺乏病的碘盐中碘的稳定性。二
A multivariate logistic regression equation to screen for diabetes - Development and validation
DIABETES CARE
IF16.6


