arrow
Return

Learning protein language contrastive models with multi-knowledge representation

delete2025-03-01
delete0
PRE
AI
徐文君 cover
徐文君 (Wenjun Xu)
S
Sun, Bifan
Z
Zihao Zhao
L
Lianggui Tang
Z
Zhou, Obo
Q
Qingyong Wang
L
Lichuan Gu *
DOI:10.1016/j.future.2024.107580delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Protein representation learning plays a crucial role in obtaining a comprehensive understanding of biological regulatory mechanisms and in developing proteins and drugs for therapeutic purposes. However, labeled proteins, such as sequenced and functionally annotated data, are incomplete and few. Thus, contrastive learning has emerged as the preferred technique for learning meaningful representations from unlabeled data samples. In addition, at present, natural proteins cannot be fully described by extracting protein knowledge from a single domain. Therefore, Pro-CoRL, a protein contrastive models framework based on multi-knowledge representation learning, was proposed in this study. In particular, Pro-CoRL smooths the objective function using convex approximation, thereby improving the stability of training. Extensive experiments on predicting protein-protein interaction types and clustering protein families have confirmed the high accuracy and robustness of Pro-CoRL.
Keywords:
Protein representation learning
Contrastive learning
Multi-knowledge embeddings
Convex approximation

Journal

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
Papers:
6.8K
Citations:
2.3W

Organization

A
Anhui Agricultural University
Scholars:
1.2W
Papers: 5.7K
Citations: 1.0W