arrow
返回

Democratizing protein language models with parameter-efficient fine-tuning

delete2024-06-20
delete5
delete
OA
AI
S
Samuel Sledzieski
M
Meghana Kshirsagar
M
Minkyung Baek
R
Rahul Dodhia
J
Juan Lavista Ferres *
B
Bonnie Berger *
DOI:10.1073/pnas.2405840121delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Proteomics has been revolutionized by large protein language models (PLMs), which learn unsupervised representations from large corpora of sequences. These models are typically fine-tuned in a supervised setting to adapt the model to specific downstream tasks. However, the computational and memory footprint of fine-tuning (FT) large PLMs presents a barrier for many research groups with limited computational resources. Natural language processing has seen a similar explosion in the size models, where these challenges have been addressed by methods for parameter -efficient fine-tuning (PEFT). In this work, we introduce this paradigm to proteomics through leveraging the parameter -efficient method LoRA and training new models for two important tasks: predicting protein-protein interactions (PPIs) and predicting the symmetry of homooligomer quaternary structures. We show that these approaches are competitive with traditional FT while requiring reduced memory and substantially fewer parameters. We additionally show that for the PPI prediction task, training only the classification head also remains competitive with full FT, using five orders of magnitude fewer parameters, and that each of these methods outperform stateof-the-art PPI prediction methods with substantially reduced compute. We further perform a comprehensive evaluation of the hyperparameter space, demonstrate that PEFT of PLMs is robust to variations in these hyperparameters, and elucidate where best practices for PEFT in proteomics differ from those in natural language processing. All our model adaptation and evaluation code is available open -source at https://github.com/microsoft/peft_proteomics. Thus, we provide a blueprint democratize the power of PLM adaptation to groups with limited computational resources.
Keyword:
protein language
parameter-efficient fine-tuning
protein-protein interactions
homooligomer symmetry
quaternary structure
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

P
Proceedings of the National Academy of Sciences of the United States of America
IF:
9.1
论文数:
10.8W
被引数:
73.5W

机构

M
Microsoft
学者数:
3.0K
论文数: 2.7K
被引数: 7
S
seoul national university (snu)
学者数:
7.2W
论文数: 6.6W
被引数: 86
引用论文

引用论文

err分享
err收藏
Opioid activities of β-casomorphinsΒ-酪啡肽的阿片活性
err1981-04-01
err0
PREAI
errVictor Brantl; Hansjörg Teschemacher; Julia Bläsig; Agnes Henschen; Friedrich Lottspeich
err分享
err收藏
Clustering coefficient and community structure of bipartite networks
err2008-12-01
err0
PREAI
errPeng Zhang; Jinliang Wang; Xiaojia Li; Menghui Li; Zengru Di; Ying Fan
err分享
err收藏
Large language models generate functional protein sequences across diverse families大型语言模型生成跨不同家族的功能蛋白质序列
err2023-01-26
err304
errOAAI
errMadani, Ali; Ben Krause, Ben; Greene, Eric R.; Subramanian, Subu; Mohr, Benjamin P.; Holton, James M.; Olmos, Jose Luis; Xiong, Caiming; Sun, Zachary Z. Z.; Socher, Richard; Fraser, James S.; Naik, Nikhil
err分享
err收藏
err分享
err收藏
Environmental and anthropogenic determinants of the spread of alien plant species: insights from South Africa’s quaternary catchments
err2018-01-15
err0
PREAI
errDilva Terzano; Ian Kotzé; Christo Marais; Silvio Cianciullo; Alessio Farcomeni; Paolo Caroli; Luca Malatesta; Fabio Attorre
err分享
err收藏
Mapping and assessing ecosystem services: Methods and practical applications
err2019-06-03
err0
errOAAI
errFernando Santos-Martín; Davide Geneletti; Benjamin Burkhard
err分享
err收藏
学者 查看更多内容