arrow
返回

Infusing clinical knowledge into language models by subword optimisation and embedding initialisation

delete2025-08-07
delete0
delete
OA
AI
A
Abul Hasan *
J
Jinge Wu
Q
Quang Nguyen
S
Salomé Andres
I
Imane Guellil
H
Huayu Zhang
A
Arlene Casey
B
Beatrice Alex
B
Bruce Guthrie
H
Honghan Wu *
DOI:10.1016/j.compbiomed.2025.110747delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
• 该研究提出了一种新颖的分词方法,利用基于领域特定概念(如药物、疾病)的词元全局表示,这些概念来自UMLS等本体或任务特定语料库。 • 在训练或推理阶段,采用词和句子级别的优化来选择最优的词元表示。 • 该研究提出了针对新词元的嵌入初始化方法,消除了语言模型预训练的需求。 • 使用K-Tokeniser构建的模型在自动临床编码任务中取得了显著的Micro F1分数提升,达到13%。该模型仅需50%的训练数据用于概念提取,且在自动编码任务中仅需不到20%的数据即可超越基线临床BERT模型。
Keyword:
Tokenisation
Language model
BERT
Clinical concept and relation extraction
ICD-9 coding classification
Phenotype identification
Document classification

期刊

Computers in Biology and Medicine 封面图
Computers in Biology and Medicine
IF:
6.3
论文数:
8.3K
被引数:
3.3W

机构

I
Institute of Health Informatics
学者数:
39
论文数: 21
被引数: 0
U
University of Edinburgh
学者数:
5.2W
论文数: 4.6W
被引数: 71
引用论文

引用论文

ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models
err2022-03-25
err87
errOAAI
errXue, Linting; Barua, Aditya; Constant, Noah; Al-Rfou, Rami; Narang, Sharan; Kale, Mihir; Roberts, Adam; Raffel, Colin
err分享
err收藏
err分享
err收藏
SemEHR: A general-purpose semantic search system to surface semantic data from clinical notes for tailored care, trial recruitment, and clinical research*
err2018-01-19
err0
errOAAI
errHonghan Wu; Giulia Toti; Katherine I Morley; Zina M Ibrahim; Amos Folarin; Richard Jackson; Ismail Kartoglu; Asha Agrawal; Clive Stringer; Darren Gale; Genevieve Gorrell; Angus Roberts; Matthew Broadbent; Robert Stewart; Richard JB Dobson
err分享
err收藏
Limitations of Transformers on Clinical Text Classification《变形金刚》对临床文本分类的局限性
err2021-09-01
err75
errOAAI
errGao, Shang; Alawad, Mohammed; Young, M. Todd; Gounley, John; Schaefferkoetter, Noah; Yoon, Hong Jun; Wu, Xiao-Cheng; Durbin, Eric B.; Doherty, Jennifer; Stroup, Antoinette; Coyle, Linda; Tourassi, Georgia
err分享
err收藏
CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison
err2019-07-17
err0
errOAAI
errJeremy Irvin; Pranav Rajpurkar; Michael Ko; Yifan Yu; Silviana Ciurea-Ilcus; Chris Chute; Henrik Marklund; Behzad Haghgoo; Robyn Ball; Katie Shpanskaya; Jayne Seekins; David A. Mong; Safwan S. Halabi; Jesse K. Sandberg; Ricky Jones; David B. Larson; Curtis P. Langlotz; Bhavik N. Patel; Matthew P. Lungren; Andrew Y. Ng
err分享
err收藏
没有更多内容