arrow
返回

Deleterious synonymous mutation identification based on selective ensemble strategy

delete2023-01-05
delete0
PRE
AI
L
Lihua Wang
T
Tao Zhang
L
Lihong Yu
郑春厚 封面图
郑春厚 (Chun-Hou Zheng)
W
Wenguang Yin *
J
Junfeng Xia *
张
张铁军 (Tiejun Zhang) *
DOI:10.1093/bib/bbac598delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Although previous studies have revealed that synonymous mutations contribute to various human diseases, distinguishing deleterious synonymous mutations from benign ones is still a challenge in medical genomics. Recently, computational tools have been introduced to predict the harmfulness of synonymous mutations. However, most of these computational tools rely on balanced training sets without considering abundant negative samples that could result in deficient performance. In this study, we propose a computational model that uses a selective ensemble to predict deleterious synonymous mutations (seDSM). We construct several candidate base classifiers for the ensemble using balanced training subsets randomly sampled from the imbalanced benchmark training sets. The diversity measures of the base classifiers are calculated by the pairwise diversity metrics, and the classifiers with the highest diversities are selected for integration using soft voting for synonymous mutation prediction. We also design two strategies for filling in missing values in the imbalanced dataset and constructing models using different pairwise diversity metrics. The experimental results show that a selective ensemble based on double fault with the ensemble strategy EKNNI for filling in missing values is the most effective scheme. Finally, using 40-dimensional biology features, we propose a novel model based on a selective ensemble for predicting deleterious synonymous mutations (seDSM). seDSM outperformed other state-of-the-art methods on the independent test sets according to multiple evaluation indicators, indicating that it has an outstanding predictive performance for deleterious synonymous mutations. We hope that seDSM will be useful for studying deleterious synonymous mutations and advancing our understanding of synonymous mutations. The source code of seDSM is freely accessible at https://github.com/xialab-ahu/seDSM.git.
Keyword:
synonymous mutation
machine learning
selective ensemble
imbalanced data

期刊

Briefings in Bioinformatics 封面图
Briefings in Bioinformatics
IF:
7.7
论文数:
5.8K
被引数:
2.7W

机构

G
Guangzhou Medical University
学者数:
2.7W
论文数: 1.4W
被引数: 3.2W
G
Guangzhou Laboratory
学者数:
1.0K
论文数: 553
被引数: 12
A
anhui university
学者数:
1.9W
论文数: 1.2W
被引数: 24
学者 查看更多机构
引用论文

引用论文

Comparison and integration of computational methods for deleterious synonymous mutation prediction
err2019-06-03
err54
PREAI
errCheng, Na; Li, Menglu; Zhao, Le; Zhang, Bo; Yang, Yuhua; Zheng, Chun-Hou; Xia, Junfeng
err分享
err收藏
A survey on ensemble learning集成学习研究综述
err2019-08-30
err1.1K
PREAI
errDong, Xibin; Yu, Zhiwen; Cao, Wenming; Shi, Yifan; Ma, Qianli
err分享
err收藏
Fluorescein labeling of Fab′ while preserving single thiol
err1988-08-01
err0
PREAI
errG.P. Der-Balian; N. Kameda; G.L. Rowley
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
学者 查看更多内容