arrow
Return

Deleterious synonymous mutation identification based on selective ensemble strategy

delete2023-01-05
delete0
PRE
AI
L
Lihua Wang
T
Tao Zhang
L
Lihong Yu
郑春厚 cover
郑春厚 (Chun-Hou Zheng)
W
Wenguang Yin *
J
Junfeng Xia *
张
张铁军 (Tiejun Zhang) *
DOI:10.1093/bib/bbac598delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Although previous studies have revealed that synonymous mutations contribute to various human diseases, distinguishing deleterious synonymous mutations from benign ones is still a challenge in medical genomics. Recently, computational tools have been introduced to predict the harmfulness of synonymous mutations. However, most of these computational tools rely on balanced training sets without considering abundant negative samples that could result in deficient performance. In this study, we propose a computational model that uses a selective ensemble to predict deleterious synonymous mutations (seDSM). We construct several candidate base classifiers for the ensemble using balanced training subsets randomly sampled from the imbalanced benchmark training sets. The diversity measures of the base classifiers are calculated by the pairwise diversity metrics, and the classifiers with the highest diversities are selected for integration using soft voting for synonymous mutation prediction. We also design two strategies for filling in missing values in the imbalanced dataset and constructing models using different pairwise diversity metrics. The experimental results show that a selective ensemble based on double fault with the ensemble strategy EKNNI for filling in missing values is the most effective scheme. Finally, using 40-dimensional biology features, we propose a novel model based on a selective ensemble for predicting deleterious synonymous mutations (seDSM). seDSM outperformed other state-of-the-art methods on the independent test sets according to multiple evaluation indicators, indicating that it has an outstanding predictive performance for deleterious synonymous mutations. We hope that seDSM will be useful for studying deleterious synonymous mutations and advancing our understanding of synonymous mutations. The source code of seDSM is freely accessible at https://github.com/xialab-ahu/seDSM.git.
Keywords:
synonymous mutation
machine learning
selective ensemble
imbalanced data

Journal

Briefings in Bioinformatics cover
Briefings in Bioinformatics
IF:
7.7
Papers:
5.8K
Citations:
2.7W

Organization

G
Guangzhou Medical University
Scholars:
2.7W
Papers: 1.4W
Citations: 3.2W
G
Guangzhou Laboratory
Scholars:
1.0K
Papers: 553
Citations: 12
A
anhui university
Scholars:
1.9W
Papers: 1.2W
Citations: 24
researcher View more organizations
Cited Papers

Cited Papers

Comparison and integration of computational methods for deleterious synonymous mutation prediction
err2019-06-03
err54
PREAI
errCheng, Na; Li, Menglu; Zhao, Le; Zhang, Bo; Yang, Yuhua; Zheng, Chun-Hou; Xia, Junfeng
errShare
errSave
A survey on ensemble learning
err2019-08-30
err1.1K
PREAI
errDong, Xibin; Yu, Zhiwen; Cao, Wenming; Shi, Yifan; Ma, Qianli
errShare
errSave
Fluorescein labeling of Fab′ while preserving single thiol
err1988-08-01
err0
PREAI
errG.P. Der-Balian; N. Kameda; G.L. Rowley
errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
Diversity measures for one-class classifier ensembles
err2014-02-01
err59
PREAI
errKrawczyk, Bartosz; Wozniak, Michal
errShare
errSave
researcher View more