返回
Data mining and life sciences applications on the grid
DOI:10.1002/widm.1090.png)
摘要
En 中文
Data mining (DM) is increasingly used in the analysis of data generated in life sciences, including biological data produced in several disciplines such as genomics and proteomics, medical data produced in clinical practice, and administrative data produced in health care. The difficulty in mining such data is twofold. First of all, data in life sciences are inherently heterogeneous, spanning from molecular level data to clinical and administrative data. Second, data in life sciences are produced at an increasing rate and data repositories are becoming very large. Thus, the management and analysis of such data is becoming a main bottleneck in biomedical research. The main goal of this paper is to review the main methodologies to mine life sciences data and the ways they are coupled to high-performance infrastructures and systems that result in an efficient analysis. This paper recalls basic concepts of DM, grids, and distributed DM on grids, and reviews main approaches to mine biomedical data on high-performance infrastructures with special focus on the analysis of genomics, proteomics, and interactomics data, and the exploration of magnetic resonance images in neurosciences. The paper can be of interest both to bioinformaticians, who can learn how to exploit high performance infrastructures to mine life sciences data, and to computer scientists, who can address the heterogeneity and the high volumes of life sciences data at the data management, algorithm, and user interface layers. (c) 2013 Wiley Periodicals, Inc.
Keyword:
SUPPORT VECTOR MACHINES
MASS-SPECTROMETRY DATA
PROTEOMICS
REPRESENTATION
ANNOTATION
ONTOLOGIES
FRAMEWORK
SERVICES
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
11.7
论文数:
548
被引数:
5.3K
机构
引用论文
Results of a phase 1, randomized, placebo-controlled first-in-human trial of griffithsin formulated in a carrageenan vaginal gel
PLOS ONE
IF0

