arrow
返回

Microarray missing data imputation based on a set theoretic framework and biological knowledge

delete2006-03-06
delete87
delete
OA
AI
X
Xiangchao Gan
L
Liew, AWC
Y
Yan, H
DOI:10.1093/nar/gkl047delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Gene expressions measured using microarrays usually suffer from the missing value problem. However, in many data analysis methods, a complete data matrix is required. Although existing missing value imputation algorithms have shown good performance to deal with missing values, they also have their limitations. For example, some algorithms have good performance only when strong local correlation exists in data while some provide the best estimate when data is dominated by global structure. In addition, these algorithms do not take into account any biological constraint in their imputation. In this paper, we propose a set theoretic framework based on projection onto convex sets (POCS) for missing data imputation. POCS allows us to incorporate different types of a priori knowledge about missing values into the estimation process. The main idea of POCS is to formulate every piece of prior knowledge into a corresponding convex set and then use a convergence-guaranteed iterative procedure to obtain a solution in the intersection of all these sets. In this work, we design several convex sets, taking into consideration the biological characteristic of the data: the first set mainly exploit the local correlation structure among genes in microarray data, while the second set captures the global correlation structure among arrays. The third set (actually a series of sets) exploits the biological phenomenon of synchronization loss in microarray experiments. In cyclic systems, synchronization loss is a common phenomenon and we construct a series of sets based on this phenomenon for our POCS imputation algorithm. Experiments show that our algorithm can achieve a significant reduction of error compared to the KNNimpute, SVDimpute and LSimpute methods.
Keyword:
GENE-EXPRESSION DATA
ARTIFACT REDUCTION
CLUSTER-ANALYSIS
CONSTRAINT SET
ALGORITHM
IDENTIFICATION
PROJECTION
PROFILE
CELLS
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Nucleic Acids Research 封面图
Nucleic Acids Research
IF:
13.1
论文数:
3.6W
被引数:
29.0W

机构

暂无机构信息
引用论文

引用论文

暂无论文信息