返回
Comparison of Feature Selection Methods for Cross-Laboratory Microarray Analysis
DOI:10.1109/TCBB.2013.70.png)
摘要
En 中文
The amount of gene expression data of microarray has grown exponentially. To apply them for extensiVe studies, integrated analysis of cross-laboratory (cross-lab) data'becomes a trend, and thus, choosing an appropriate feature selection method is an essential issue. This paper focuses on,feature selection for Affymetrix.(Affy) microarray studies across differentlabs. We investigate four feature selection methods: t-test, significance analysis of microarrays (SAM), rank products'(RP), and random forest (RF)..The four methods are applied'to acute lymphoblastic leukemia, acute myeloid leukemia, breast cancer, and lung cancer Affy datawhich consist of three cross-lab data sets each. We utilize a rank-based normalization method to reduce thetias from cross-lab-data sets. Training on one data set or two combined data sets to test the remaining data set(s) are both' considered, Balanced accuracy is used for prediction evaluation. This study provides comprehensive comparisons of the four feature selection methods in cross-lab microarray analysis. Results show, that SAM has the best classification performance. RF also getS high classification accuracy, but it is not as stable as SAM. The most naive method is t-test, but its performance is the worst among the four methods. In thiS study, we further discuss the influence from the number of training samples, the number of selected genes, and the issue of unbalanced data sets.
Keyword:
Microarray data analysis
feature selection
cancer
cross-laboratory experiment
期刊
I
IF:
3.4
论文数:
3.3K
被引数:
6.4K
机构
引用论文
Gene Expression Omnibus: NCBI gene expression and hybridization array data repository基因表达综合: NCBI基因表达和杂交阵列数据库
NUCLEIC ACIDS RESEARCH
IF13.1
The MicroArray Quality Control (MAQC)-IIII study of common practices for the development and validation of microarray-based predictive models微阵列质量控制 (MAQC)-IIII基于微阵列的预测模型的开发和验证的通用实践研究
NATURE BIOTECHNOLOGY
IF41.7
An expression-based site of origin diagnostic method designed for clinical application to cancer of unknown origin
CANCER RESEARCH
IF16.6
Tackling the widespread and critical impact of batch effects in high-throughput data解决高通量数据中批量效应的广泛和关键影响

