Return
Learning protein language contrastive models with multi-knowledge representation
DOI:10.1016/j.future.2024.107580.png)
Abstract
En 中文
Protein representation learning plays a crucial role in obtaining a comprehensive understanding of biological regulatory mechanisms and in developing proteins and drugs for therapeutic purposes. However, labeled proteins, such as sequenced and functionally annotated data, are incomplete and few. Thus, contrastive learning has emerged as the preferred technique for learning meaningful representations from unlabeled data samples. In addition, at present, natural proteins cannot be fully described by extracting protein knowledge from a single domain. Therefore, Pro-CoRL, a protein contrastive models framework based on multi-knowledge representation learning, was proposed in this study. In particular, Pro-CoRL smooths the objective function using convex approximation, thereby improving the stability of training. Extensive experiments on predicting protein-protein interaction types and clustering protein families have confirmed the high accuracy and robustness of Pro-CoRL.
Keywords:
Protein representation learning
Contrastive learning
Multi-knowledge embeddings
Convex approximation
Journal
F
IF:
6.1
Papers:
6.8K
Citations:
2.3W

