Return
Infusing clinical knowledge into language models by subword optimisation and embedding initialisation
DOI:10.1016/j.compbiomed.2025.110747.png)
Abstract
En 中文
• The study proposes a novel tokenisation method utilising global representations of tokens based on domain-specific concepts (e.g., drugs, diseases) from ontologies like UMLS or task-specific corpora. • At training or inference, word and sentence-level optimisation is used to select the optimal token representations. • It proposes an embedding initialisation approach for new tokens, eliminating the need for pre-training the language models. • The Model built using K-Tokeniser achieves a notable 13% increase in Micro F1 score for automated clinical coding. It requires only 50% of training data for concept extraction and less than 20% for automated coding to outperform the baseline clinicalBERT model.
Keywords:
Tokenisation
Language model
BERT
Clinical concept and relation extraction
ICD-9 coding classification
Phenotype identification
Document classification
Journal
IF:
6.3
Papers:
8.3K
Citations:
3.3W

