arrow
返回

Transcriptomics and epigenetic data integration learning module on Google Cloud

delete2024-08-05
delete1
delete
OA
AI
N
Nathan A. Ruprecht
J
Joshua D Kennedy
B
Benu Bansal
S
Sonalika Singhal
D
Donald A. Sens
A
Angela Maggio
V
Valena Doe
D
Dale Hawkins
R
Ross Campbel
K
Kyle A. O’Connell
J
Jappreet Singh Gill
K
Kalli Schaefer
S
Sandeep K. Singhal *
DOI:10.1093/bib/bbae352delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Multi-omics (genomics, transcriptomics, epigenomics, proteomics, metabolomics, etc.) research approaches are vital for understanding the hierarchical complexity of human biology and have proven to be extremely valuable in cancer research and precision medicine. Emerging scientific advances in recent years have made high-throughput genome-wide sequencing a central focus in molecular research by allowing for the collective analysis of various kinds of molecular biological data from different types of specimens in a single tissue or even at the level of a single cell. Additionally, with the help of improved computational resources and data mining, researchers are able to integrate data from different multi-omics regimes to identify new prognostic, diagnostic, or predictive biomarkers, uncover novel therapeutic targets, and develop more personalized treatment protocols for patients. For the research community to parse the scientifically and clinically meaningful information out of all the biological data being generated each day more efficiently with less wasted resources, being familiar with and comfortable using advanced analytical tools, such as Google Cloud Platform becomes imperative. This project is an interdisciplinary, cross-organizational effort to provide a guided learning module for integrating transcriptomics and epigenetics data analysis protocols into a comprehensive analysis pipeline for users to implement in their own work, utilizing the cloud computing infrastructure on Google Cloud. The learning module consists of three submodules that guide the user through tutorial examples that illustrate the analysis of RNA-sequence and Reduced-Representation Bisulfite Sequencing data. The examples are in the form of breast cancer case studies, and the data sets were procured from the public repository Gene Expression Omnibus. The first submodule is devoted to transcriptomics analysis with the RNA sequencing data, the second submodule focuses on epigenetics analysis using the DNA methylation data, and the third submodule integrates the two methods for a deeper biological understanding. The modules begin with data collection and preprocessing, with further downstream analysis performed in a Vertex AI Jupyter notebook instance with an R kernel. Analysis results are returned to Google Cloud buckets for storage and visualization, removing the computational strain from local resources. The final product is a start-to-finish tutorial for the researchers with limited experience in multi-omics to integrate transcriptomics and epigenetics data analysis into a comprehensive pipeline to perform their own biological research. This manuscript describes the development of a resource module that is part of a learning platform named ``NIGMS Sandbox for Cloud-based Learning'' https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox [] at the beginning of this Supplement. This module delivers learning materials on the analysis of bulk and single-cell ATAC-seq data in an interactive format that uses appropriate cloud resources for data access and analyses. Multi-omics (genomics, transcriptomics, epigenomics, proteomics, metabolomics, etc.) research approaches are vital for understanding the hierarchical complexity of human biology and have proven to be extremely valuable in cancer research and precision medicine. Emerging scientific advances in recent years have made high-throughput genome-wide sequencing a central focus in molecular research by allowing for the collective analysis of various kinds of molecular biological data from different types of specimens in a single tissue or even at the level of a single cell. Additionally, with the help of improved computational resources and data mining, researchers are able to integrate data from different multi-omics regimes to identify new prognostic, diagnostic, or predictive biomarkers, uncover novel therapeutic targets, and develop more personalized treatment protocols for patients. For the research community to parse the scientifically and clinically meaningful information out of all the biological data being generated each day more efficiently with less wasted resources, being familiar with and comfortable using advanced analytical tools, such as Google Cloud Platform becomes imperative. This project is an interdisciplinary, cross-organizational effort to provide a guided learning module for integrating transcriptomics and epigenetics data analysis protocols into a comprehensive analysis pipeline for users to implement in their own work, utilizing the cloud computing infrastructure on Google Cloud. The learning module consists of three submodules that guide the user through tutorial examples that illustrate the analysis of RNA-sequence and Reduced-Representation Bisulfite Sequencing data. The examples are in the form of breast cancer case studies, and the data sets were procured from the public repository Gene Expression Omnibus. The first submodule is devoted to transcriptomics analysis with the RNA sequencing data, the second submodule focuses on epigenetics analysis using the DNA methylation data, and the third submodule integrates the two methods for a deeper biological understanding. The modules begin with data collection and preprocessing, with further downstream analysis performed in a Vertex AI Jupyter notebook instance with an R kernel. Analysis results are returned to Google Cloud buckets for storage and visualization, removing the computational strain from local resources. The final product is a start-to-finish tutorial for the researchers with limited experience in multi-omics to integrate transcriptomics and epigenetics data analysis into a comprehensive pipeline to perform their own biological research. This manuscript describes the development of a resource module that is part of a learning platform named ``NIGMS Sandbox for Cloud-based Learning'' https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox [] at the beginning of this Supplement. This module delivers learning materials on the analysis of bulk and single-cell ATAC-seq data in an interactive format that uses appropriate cloud resources for data access and analyses.
Keyword:
transcriptomics
epigenomics
Google Cloud computing
multi-omics integration
DNA methylation
R Bioconductor
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Briefings in Bioinformatics 封面图
Briefings in Bioinformatics
IF:
7.7
论文数:
5.8K
被引数:
2.7W

机构

U
university of north dakota grand forks
学者数:
3.3K
论文数: 2.7K
被引数: 3
D
deloitte touche tohmatsu limited
学者数:
406
论文数: 271
被引数: 0
G
Google Incorporated
学者数:
3.5K
论文数: 1.8K
被引数: 8
学者 查看更多机构
引用论文

引用论文

Multi-omics analyses identify molecular signatures with prognostic values in different heart failure aetiologies
err2023-02-01
err13
errOAAI
errAboumsallem, Joseph Pierre; Shi, Canxia; De Wit, Sanne; Markousis-Mavrogenis, George; Bracun, Valentina; Eijgenraam, Tim R.; Hoes, Martijn F.; Meijers, Wouter C.; Screever, Elles M.; Schouten, Marloes E.; Voors, Adriaan A.; Sillje, Herman H. W.; De Boer, Rudolf A.
err分享
err收藏
Multi-omics analysis of tumor angiogenesis characteristics and potential epigenetic regulation mechanisms in renal clear cell carcinoma
err2021-03-24
err37
errOAAI
errZheng, Wenzhong; Zhang, Shiqiang; Guo, Huan; Chen, Xiaobao; Huang, Zhangcheng; Jiang, Shaoqin; Li, Mengqiang
err分享
err收藏
Multi-omics data integration considerations and study design for biological systems and disease
err2021-01-01
err107
errOAAI
errGraw, Stefan; Chappell, Kevin; Washam, Charity L.; Gies, Allen; Bird, Jordan; Robeson, Michael S., II; Byrum, Stephanie D.
err分享
err收藏
Lung Cancer Screening Among U.S. Military Veterans by Health Status and Race and Ethnicity, 2017–2020: A Cross-Sectional Population-Based Study
err2023-06-01
err0
errOAAI
errAlison S. Rustagi; Amy L. Byers; James K. Brown; Natalie Purcell; Christopher G. Slatore; Salomeh Keyhani
err分享
err收藏
err分享
err收藏
Smoking induces coordinated DNA methylation and gene expression changes in adipose tissue with consequences for metabolic health
err2018-10-20
err99
errOAAI
errTsai, Pei-Chien; Glastonbury, Craig A.; Eliot, Melissa N.; Bollepalli, Sailalitha; Yet, Idil; Castillo-Fernandez, Juan E.; Carnero-Montoro, Elena; Hardiman, Thomas; Martin, Tiphaine C.; Vickers, Alice; Mangino, Massimo; Ward, Kirsten; Pietilaeinen, Kirsi H.; Deloukas, Panos; Spector, Tim D.; Vinuela, Ana; Loucks, Eric B.; Ollikainen, Miina; Kelsey, Karl T.; Small, Kerrin S.; Bell, Jordana T.
err分享
err收藏
学者 查看更多内容