返回
Gene function finding through cross-organism ensemble learning
DOI:10.1186/s13040-021-00239-w.png)
摘要
En 中文
Background Structured biological information about genes and proteins is a valuable resource to improve discovery and understanding of complex biological processes via machine learning algorithms. Gene Ontology (GO) controlled annotations describe, in a structured form, features and functions of genes and proteins of many organisms. However, such valuable annotations are not always reliable and sometimes are incomplete, especially for rarely studied organisms. Here, we present GeFF (Gene Function Finder), a novel cross-organism ensemble learning method able to reliably predict new GO annotations of a target organism from GO annotations of another source organism evolutionarily related and better studied. Results Using a supervised method, GeFF predicts unknown annotations from random perturbations of existing annotations. The perturbation consists in randomly deleting a fraction of known annotations in order to produce a reduced annotation set. The key idea is to train a supervised machine learning algorithm with the reduced annotation set to predict, namely to rebuild, the original annotations. The resulting prediction model, in addition to accurately rebuilding the original known annotations for an organism from their perturbed version, also effectively predicts new unknown annotations for the organism. Moreover, the prediction model is also able to discover new unknown annotations in different target organisms without retraining.We combined our novel method with different ensemble learning approaches and compared them to each other and to an equivalent single model technique. We tested the method with five different organisms using their GO annotations: Homo sapiens, Mus musculus, Bos taurus, Gallus gallus and Dictyostelium discoideum. The outcomes demonstrate the effectiveness of the cross-organism ensemble approach, which can be customized with a trade-off between the desired number of predicted new annotations and their precision.A Web application to browse both input annotations used and predicted ones, choosing the ensemble prediction method to use, is publicly available at . Conclusions Our novel cross-organism ensemble learning method provides reliable predicted novel gene annotations, i.e., functions, ranked according to an associated likelihood value. They are very valuable both to speed the annotation curation, focusing it on the prioritized new annotations predicted, and to complement known annotations available.
Keyword:
Biomolecular annotation prediction
Knowledge discovery
Ensemble learning
Transfer learning
Data representation
Gene ontology
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.1
论文数:
700
被引数:
1.5K
机构
引用论文
Bioinformatics enrichment tools: paths toward the comprehensive functional analysis of large gene lists生物信息学富集工具: 通往大型基因列表综合功能分析的路径
NUCLEIC ACIDS RESEARCH
IF13.1
Associating genes with gene ontology codes using a maximum entropy analysis of biomedical literature
GENOME RESEARCH
IF5.5
Rare earth element (REE) enrichment of the late Ediacaran Kalyus Beds (East European Platform) through diagenetic uptake
Geochemistry
IF0
IMP: a multi-species functional genomics portal for integration, visualization and prediction of protein functions and networks
NUCLEIC ACIDS RESEARCH
IF13.1
Elements Doping Strategy for Improving the Thermoelectric Properties of g-C3N4/SWCNT Composite Filmsg-C3N4/SWCNT复合薄膜热电性能提升的元素掺杂策略
Paleomagnetism of the Late Cretaceous ignimbrite from the Okhotsk-Chukotka Volcanic Belt, Kolyma-Omolon Composite Terrane: Tectonic implications晚白垩世伊格尼姆岩的古地磁研究:来自鄂霍次克-楚科奇火山带,科雷马-奥莫隆复合地体的构造意义

