arrow
Return

A self-supervised deep learning method for data-efficient training in genomics

delete2023-09-11
delete4
delete
OA
AI
H
Hüseyin Anil Gündüz
M
Martin Binder
X
Xiao-Yin To
R
René Mreches
B
Bernd Bischl
A
Alice C. McHardy
P
Philipp C. Münch *
M
Mina Rezaei *
DOI:10.1038/s42003-023-05310-2delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Deep learning in bioinformatics is often limited to problems where extensive amounts of labeled data are available for supervised classification. By exploiting unlabeled data, self-supervised learning techniques can improve the performance of machine learning models in the presence of limited labeled data. Although many self-supervised learning methods have been suggested before, they have failed to exploit the unique characteristics of genomic data. Therefore, we introduce Self-GenomeNet, a self-supervised learning technique that is custom-tailored for genomic data. Self-GenomeNet leverages reverse-complement sequences and effectively learns short- and long-term dependencies by predicting targets of different lengths. Self-GenomeNet performs better than other self-supervised methods in data-scarce genomic tasks and outperforms standard supervised training with similar to 10 times fewer labeled training data. Furthermore, the learned representations generalize well to new datasets and tasks. These findings suggest that Self-GenomeNet is well suited for large-scale, unlabeled genomic datasets and could substantially improve the performance of genomic models.
Keywords:
REPRESENTATIONS
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Communications Biology cover
Communications Biology
IF:
5.1
Papers:
1.0W
Citations:
3.2W

Organization

U
University of Munich
Scholars:
5.7W
Papers: 4.2W
Citations: 68
H
helmholtz-center for infection research
Scholars:
2.9K
Papers: 2.0K
Citations: 4