arrow
Return

Infusing clinical knowledge into language models by subword optimisation and embedding initialisation

delete2025-08-07
delete0
delete
OA
AI
A
Abul Hasan *
J
Jinge Wu
Q
Quang Nguyen
S
Salomé Andres
I
Imane Guellil
H
Huayu Zhang
A
Arlene Casey
B
Beatrice Alex
B
Bruce Guthrie
H
Honghan Wu *
DOI:10.1016/j.compbiomed.2025.110747delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• The study proposes a novel tokenisation method utilising global representations of tokens based on domain-specific concepts (e.g., drugs, diseases) from ontologies like UMLS or task-specific corpora. • At training or inference, word and sentence-level optimisation is used to select the optimal token representations. • It proposes an embedding initialisation approach for new tokens, eliminating the need for pre-training the language models. • The Model built using K-Tokeniser achieves a notable 13% increase in Micro F1 score for automated clinical coding. It requires only 50% of training data for concept extraction and less than 20% for automated coding to outperform the baseline clinicalBERT model.
Keywords:
Tokenisation
Language model
BERT
Clinical concept and relation extraction
ICD-9 coding classification
Phenotype identification
Document classification

Journal

Computers in Biology and Medicine cover
Computers in Biology and Medicine
IF:
6.3
Papers:
8.3K
Citations:
3.3W

Organization

I
Institute of Health Informatics
Scholars:
37
Papers: 20
Citations: 0
U
University of Edinburgh
Scholars:
5.2W
Papers: 4.6W
Citations: 71