arrow
Return

Automated epilepsy and seizure type phenotyping with transformer-based language models

delete2026-08-11
delete0
delete
OA
AI
E
Ellie Chang
K
Kevin Xie
D
Daniel J. Zhou
J
Jacob Korzun
E
Erin C. Conrad
D
Dan Roth
B
Brian Litt
C
Colin A. Ellis *
DOI:10.1038/s41746-026-03103-5delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Epilepsy is a common neurologic disorder characterized by recurrent, unprovoked seizures, and detailed phenotypes are essential for treatment selection, prognostication, and outcomes research. Although electronic health records provide rich longitudinal data, structured fields such as diagnostic codes fail to capture these phenotypes, which remain buried in free-text clinical notes. We developed and evaluated two transformer-based language models for automated epilepsy and seizure type phenotyping at an academic epilepsy center: a fine-tuned masked language model (BERT) and a reasoning-optimized large language model (DeepSeek-R1), benchmarked against pairwise inter-rater agreement among board-certified epileptologists. For coarser tasks, both models matched expert agreement: classifying epilepsy type as focal, generalized, or other (Matthews correlation coefficient: DeepSeek = 0.85, BERT = 0.73, human = 0.77) and seizure type as convulsive or non-convulsive (DeepSeek = 0.74, BERT = 0.60, human = 0.49). For more granular tasks, DeepSeek maintained expert-level performance, whereas BERT declined. Deploying DeepSeek-R1 on 77,049 notes from 18,566 patients yielded patient-level phenotypes consistent with expected clinical patterns, including diagnostic stabilization, seizure type co-occurrence, and outcome differences by epilepsy type. Transformer-based models can extract expert-level epilepsy phenotypes at health-system scale, demonstrating the feasibility of automated phenotyping for large-scale epilepsy research and, with prospective validation, for future clinical decision support.

Journal

npj Digital Medicine cover
npj Digital Medicine
IF:
15.1
Papers:
3.1K
Citations:
1.5W

Organization

U
University of Pennsylvania
Scholars:
1.0W
Papers: 3.7K
Citations: 11.8W