返回
Experiences with workflows for automating data-intensive bioinformatics
DOI:10.1186/s13062-015-0071-8.png)
摘要
En 中文
High-throughput technologies, such as next-generation sequencing, have turned molecular biology into a data-intensive discipline, requiring bioinformaticians to use high-performance computing resources and carry out data management and analysis tasks on large scale. Workflow systems can be useful to simplify construction of analysis pipelines that automate tasks, support reproducibility and provide measures for fault-tolerance. However, workflow systems can incur significant development and administration overhead so bioinformatics pipelines are often still built without them. We present the experiences with workflows and workflow systems within the bioinformatics community participating in a series of hackathons and workshops of the EU COST action SeqAhead. The organizations are working on similar problems, but we have addressed them with different strategies and solutions. This fragmentation of efforts is inefficient and leads to redundant and incompatible solutions. Based on our experiences we define a set of recommendations for future systems to enable efficient yet simple bioinformatics workflow construction and execution.
Keyword:
Workflow
Automation
Data-intensive
High-performance computing
Big data
Reproducibility
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
4.9
论文数:
1.4K
被引数:
2.7K
机构
引用论文
Lessons learned from implementing a national infrastructure in Sweden for storage and analysis of next-generation sequencing data
GIGASCIENCE
IF3.9

