Return
Prediction of antibody non-specificity using protein language models and biophysical parameters
L
L
S
P
A
N
M
D
DOI:10.1080/19420862.2026.2678000.png)
Abstract
En 中文
The development of therapeutic antibodies requires optimizing target binding affinity and pharmacodynamics, while ensuring high developability potential, including minimizing non-specific binding. In this study, we address this problem by predicting antibody non-specificity by two complementary approaches: (1) antibody sequence embeddings by protein language models (PLMs) and (2) a comprehensive set of sequence-based biophysical descriptors. We benchmark the PLM embeddings against interpretable, sequence-derived biophysical descriptors and use fragment-specific models (variable heavy (VH) and variable light (VL) chains, concatenated regions, and individual complementary-determining regions (CDRs) to identify region-level determinants of non-specificity in functional antibodies that bind defined targets. These models were trained on previously published human and mouse antibody data and tested on three public datasets. We show that non-specificity is best predicted from the VH domain and heavy-chain CDRs. These region-level analyses show that heavy-chain features, especially H-CDR3, dominate the predictive signal. The top performing PLM, a VH domain-based Evolutionary Scale Modeling 1 v LogisticReg model, resulted in 10-fold cross-validation accuracy of up to 71%. While predictive accuracy is comparable to classical machine-learning baselines, PLM-based embeddings provide sequence-context representations that yield consistent region-level attribution and provide complementarity to the classical models by enhancing the reliability of predicted non‑specificity scores when used in combination. Our biophysical descriptor-based analysis identified the isoelectric point as a key driver of non-specificity, consistent with previous reports. Our findings highlight the importance of biophysical properties in predicting antibody non-specificity and highlight the potential of PLMs for the development of antibody-based therapeutics. These conclusions are robust to alternative class definitions and are not driven by VH-VL mutational bias. We illustrate the generalizability and practical use of the PLM approach by extending it to therapeutic antibodies and nanobodies, providing a tool for early-stage developability assessment, and we make the models publicly available.
Keywords:
Therapeutic antibodies
non-specificity
protein-language models
machine learning
isoelectric point
Journal
IF:
7.3
Papers:
1.8K
Citations:
7.2K
