1
Return

Advancing bioinformatics with language models: components, applications, and perspectives

delete2026-07-10
delete0
delete
OA
AI
J
Jiajia Liu *
M
Mengyuan Yang
Y
Yankai Yu
H
Haixia Xu
T
Tiangang Wang
X
Xiaobo Zhou
DOI:10.1093/bib/bbag367delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large language models (LLMs) are deep learning-based artificial intelligence models that have achieved remarkable success in natural language processing. Typically composed of neural networks with billions of parameters, they are trained on massive unlabeled datasets using self-supervised or semi-supervised learning. Beyond language, LLMs hold immense potential for addressing complex bioinformatics challenges. This review provides a comprehensive overview of transformer-based model applications in genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis. We discuss critical components, including tokenization strategies for diverse biological data, transformer architectures, attention mechanisms, and pretraining approaches. We also survey currently available foundation models and their downstream applications across bioinformatics domains. Finally, we highlight major challenges that remain insufficiently addressed in prior reviews and outline future perspectives and design principles for next-generation biological language models, offering practical guidance for both users and developers.

Journal

Briefings in Bioinformatics cover
Briefings in Bioinformatics
IF:
7.7
Papers:
5.6K
Citations:
2.7W

Organization

X
Xi'an Jiaotong University Health Science Center
Scholars:
66
Papers: 17
Citations: 0
S
sichuan university
Scholars:
11.5W
Papers: 7.6W
Citations: 100
S
southwest jiaotong university
Scholars:
7.6K
Papers: 2.7K
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers