arrow
Return

Prediction of Enzyme function using interpretable optimized Ensemble learning framework

delete2025-09-01
delete0
delete
OA
AI
S
Saikat Dhibar
S
Sumon Basak
B
Biman Jana
DOI:10.1039/D5SC04513Ddelete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Accurate prediction of enzyme function; particularly for newly discovered uncharacterized sequences; is immensely important for modern biology research. Recently machine learning (ML) based methods have shown promises. However; such tools often suffer from complexity in feature extraction; interpretability; and generalization ability. In this study; we construct the dataset for enzyme functions and present an interpretable ML method; SOLVE (Soft-Voting Optimized Learning for Versatile Enzymes) that addresses these issues by using only combination of tokenized subsequences from the protein's primary sequence for classification. SOLVE utilizes an ensemble learning framework integrating random forest (RF); light gradient boosting machine (LightGBM) and decision tree (DT) models with an optimized weighted strategy which enhances prediction accuracy; distinguishes enzymes from non-enzymes; and predicts enzyme commission (EC) numbers for mono- and multi-functional enzymes. The focal loss penalty in SOLVE effectively mitigates class imbalance; refining functional annotation accuracy. Additionally; SOLVE provides interpretability through Shapley analyses; identifying functional motifs at catalytic and allosteric sites of enzymes. By leveraging only primary sequence data; SOLVE streamlines high-throughput enzyme function prediction for functionally uncharacterized sequences and outperforms existing tools across all evaluation metrics on independent datasets. With high prediction accuracy and its identification ability of functional regions; SOLVE can become a promising tool in different fields of biology and therapeutic drug design.
Keywords:
enzyme function prediction
machine learning
interpretable models
enzyme commission (EC) numbers
class imbalance mitigation
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Chemical Science cover
Chemical Science
IF:
7.4
Papers:
1.7W
Citations:
9.3W

Organization

No organization information available
Cited Papers

Cited Papers

CD-HIT: accelerated for clustering the next-generation sequencing data
err2012-10-11
err0
errOAAI
errLimin Fu; Beifang Niu; Zhengwei Zhu; Sitao Wu; Weizhong Li
errShare
errSave
Biocatalysis: Enzymatic Synthesis for Industrial Applications
err2020-08-17
err0
errOAAI
errShuke Wu; Radka Snajdrova; Jeffrey C. Moore; Kai Baldenius; Uwe T. Bornscheuer
errShare
errSave
Important therapeutic targets in chronic myelogenous leukemia
err2007-02-22
err96
errOAAI
errKantarjian, Hagop M.; Giles, Francis; Quintas-Cardama, Alfonso; Cortes, Jorge
errShare
errSave
Motif-Based Protein Sequence Classification Using Neural Networks
err2005-02-01
err0
errOAAI
errKonstantinos Blekas; Dimitrios I. Fotiadis; Aristidis Likas
errShare
errSave
DeEPn: a deep neural network based tool for enzyme functional annotation
err2020-04-22
err0
PREAI
errRahul Semwal; Imlimaong Aier; Pankaj Tyagi; Pritish Kumar Varadwaj
errShare
errSave
DEEPre: sequence-based enzyme EC number prediction by deep learning
err2017-10-23
err0
errOAAI
errYu Li; Sheng Wang; Ramzan Umarov; Bingqing Xie; Ming Fan; Lihua Li; Xin Gao
errShare
errSave
researcher View more