arrow
Return

Pre-training phenotyping classifiers

delete2021-01-01
delete2
delete
OA
AI
D
Dmitriy Dligach *
M
Majid Afshar
T
Timothy A. Miller
DOI:10.1016/j.jbi.2020.103626delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Recent transformer-based pre-trained language models have become a de facto standard for many text classification tasks. Nevertheless, their utility in the clinical domain, where classification is often performed at encounter or patient level, is still uncertain due to the limitation on the maximum length of input. In this work, we introduce a self-supervised method for pre-training that relies on a masked token objective and is free from the limitation on the maximum input length. We compare the proposed method with supervised pre-training that uses billing codes as a source of supervision. We evaluate the proposed method on one publicly-available and three in-house datasets using the standard evaluation metrics such as the area under the ROC curve and F1 score. We find that, surprisingly, even though self-supervised pre-training performs slightly worse than supervised, it still preserves most of the gains from pre-training.
Keywords:
Natural language processing
Automatic phenotyping
Pre-training
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Biomedical Informatics cover
Journal of Biomedical Informatics
IF:
4.5
Papers:
3.5K
Citations:
1.9W

Organization

L
Loyola University Chicago
Scholars:
8.0K
Papers: 6.0K
Citations: 5.8K
U
university of wisconsin madison
Scholars:
3.8W
Papers: 2.9W
Citations: 53
University of Wisconsin System cover
University of Wisconsin System
Scholars:
6.7W
Papers: 5.8W
Citations: 382
researcher View more organizations