arrow
返回

Dealing with missing values in large-scale studies: microarray data imputation and beyond

delete2009-12-04
delete147
delete
OA
AI
T
Tero Aittokallio *
DOI:10.1093/bib/bbp059delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
High-throughput biotechnologies, such as gene expression microarrays or mass-spectrometry-based proteomic assays, suffer from frequent missing values due to various experimental reasons. Since the missing data points can hinder downstream analyses, there exists a wide variety of ways in which to deal with missing values in large-scale data sets. Nowadays, it has become routine to estimate (or impute) the missing values prior to the actual data analysis. After nearly a decade since the publication of the first missing value imputation methods for gene expression microarray data, new imputation approaches are still being developed at an increasing rate. However, what is lagging behind is a systematic and objective evaluation of the strengths and weaknesses of the different approaches when faced with different types of data sets and experimental questions. In this review, the present strategies for missing value imputation and the measures for evaluating their performance are described. The imputation methods are first reviewed in the context of gene expression microarray data, since most of the methods have been developed for estimating gene expression levels; then, we turn to other large-scale data sets that also suffer from the problems posed by missing values, together with pointers to possible imputation approaches in these settings. Along with a description of the basic principles behind the different imputation approaches, the review tries to provide practical guidance for the users of high-throughput technologies on how to choose the imputation tool for their data and questions, and some additional research directions for the developers of imputation methodologies.
Keyword:
missing value imputation
gene expression microarrays
mass-spectrometry proteomics
statistical modelling
biomarker discovery
disease classification

期刊

Briefings in Bioinformatics 封面图
Briefings in Bioinformatics
IF:
7.7
论文数:
5.8K
被引数:
2.7W

机构

暂无机构信息
引用论文

引用论文

Synthesis of Placental Protein 12 by Human Decidua
err1985-04-01
err0
PREAI
errEEVA-MARJA RUTANEN; RIITTA KOISTINEN; TORSTEN WAHLSTROM; HANS BOHN; TAPIO RANTA; MARKKU SEPPALA
err分享
err收藏
Ameliorative missing value imputation for robust biological knowledge inference
err2008-08-01
err20
PREAI
errSehgal, Muhammad Shoaib B.; Gondal, Iqbal; Dooley, Laurence S.; Coppel, Ross
err分享
err收藏
err分享
err收藏
err分享
err收藏
Developing and Validating a Model for a Plant Growth Regulator
err1995-11-01
err0
PREAI
errK. Raja Reddy; Mariquita L. Boone; A. R. Reddy; Harry F. Hodges; Sammy B. Turner; James M. McKinion
err分享
err收藏
学者 查看更多内容