arrow
Return

Content Coding of Psychotherapy Transcripts Using Labeled Topic Models

delete2017-03-01
delete32
delete
OA
AI
G
Garren Gaut *
M
Mark Steyvers
Z
Zac E. Imel
D
David C. Atkins
P
Padhraic Smyth
DOI:10.1109/JBHI.2015.2503985delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Psychotherapy represents a broad class of medical interventions received by millions of patients each year. Unlike most medical treatments, its primary mechanisms are linguistic; i.e., the treatment relies directly on a conversation between a patient and provider. However, the evaluation of patient-provider conversation suffers from critical shortcomings, including intensive labor requirements, coder error, nonstandardized coding systems, and inability to scale up to larger data sets. To overcome these shortcomings, psychotherapy analysis needs a reliable and scalable method for summarizing the content of treatment encounters. We used a publicly available psychotherapy corpus from Alexander Street press comprising a large collection of transcripts of patient-provider conversations to compare coding performance for two machine learning methods. We used the labeled latent Dirichlet allocation (LLDA) model to learn associations between text and codes, to predict codes in psychotherapy sessions, and to localize specific passages of within-session text representative of a session code. We compared the L-LDA model to a baseline lasso regression model using predictive accuracy and model generalizability (measured by calculating the area under the curve (AUC) from the receiver operating characteristic curve). The L-LDA model outperforms the lasso logistic regression model at predicting session-level codes with average AUC scores of 0.79, and 0.70, respectively. For fine-grained level coding, L-LDA and logistic regression are able to identify specific talk-turns representative of symptom codes. However, model performance for talk-turn identification is not yet as reliable as human coders. We conclude that the L-LDA model has the potential to be an objective, scalable method for accurate automated coding of psychotherapy sessions that perform better than comparable discriminative methods at session-level coding and can also predict fine-grained codes.
Keywords:
Clinical communication
conversation analysis
labeled latent Dirichlet allocation (L-LDA)
machine learning
multilabel document classification

Journal

IEEE Journal of Biomedical and Health Informatics cover
IEEE Journal of Biomedical and Health Informatics
IF:
6.8
Papers:
4.5K
Citations:
2.0W

Organization

U
University of Utah
Scholars:
2.9W
Papers: 2.2W
Citations: 4.6W
U
Utah System of Higher Education
Scholars:
4.6W
Papers: 4.0W
Citations: 161
University of California System cover
University of California System
Scholars:
37.5W
Papers: 33.7W
Citations: 6.6K
U
university of california irvine
Scholars:
2.3W
Papers: 1.7W
Citations: 55
researcher View more organizations