返回
Retrieval-Based Diagnostic Decision Support: Mixed Methods Study
DOI:10.2196/50209.png)
摘要
En 中文
Background: Diagnostic errors pose significant health risks and contribute to patient mortality. With the growing accessibility of electronic health records, machine learning models offer a promising avenue for enhancing diagnosis quality. Current research has primarily focused on a limited set of diseases with ample training data, neglecting diagnostic scenarios with limited data availability. Objective: This study aims to develop an information retrieval (IR)-based framework that accommodates data sparsity to facilitate broader diagnostic decision support. Methods: We introduced an IR-based diagnostic decision support framework called CliniqIR. It uses clinical text records, the Unified Medical Language System Metathesaurus, and 33 million PubMed abstracts to classify a broad spectrum of diagnoses independent of training data availability. CliniqIR is designed to be compatible with any IR framework. Therefore, we implemented it using both dense and sparse retrieval approaches. We compared CliniqIR's performance to that of pretrained clinical transformer models such as Clinical Bidirectional Encoder Representations from Transformers (ClinicalBERT) in supervised and zero-shot settings. Subsequently, we combined the strength of supervised fine-tuned ClinicalBERT and CliniqIR to build an ensemble Results: On a complex diagnosis data set (DC3) without any training data, CliniqIR models returned the correct diagnosis within their top 3 predictions. On the Medical Information Mart for Intensive Care III data set, CliniqIR models surpassed ClinicalBERT in predicting diagnoses with <5 training samples by an average difference in mean reciprocal rank of 0.10. In a zero-shot setting where models received no disease-specific training, CliniqIR still outperformed the pretrained transformer models with a greater mean reciprocal rank of at least 0.10. Furthermore, in most conditions, our ensemble framework surpassed the performance of its individual components, demonstrating its enhanced ability to make precise diagnostic predictions. Conclusions: Our experiments highlight the importance of IR in leveraging unstructured knowledge resources to identify infrequently encountered diagnoses. In addition, our ensemble framework benefits from combining the complementary strengths of the supervised and retrieval-based models to diagnose a broad spectrum of diseases.
Keyword:
clinical decision support
rare diseases
ensemble learning
retrieval-augmented learning
machine learning
electronic health records
natural language processing
retrieval augmented generation
RAG
electronic health record
EHR
data sparsity
information retrieval
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.8
论文数:
1.5K
被引数:
4.3K
机构
引用论文
Comparative Accuracy of Diagnosis by Collective Intelligence of Multiple Physicians vs Individual Physicians
JAMA NETWORK OPEN
IF9.7
Inverse association between intelligence quotient and urinary retinol binding protein in Chinese school-age children with low blood lead levels: Results from a cross-sectional investigation
Chemosphere
IF0
Comparison of Neural Language Modeling Pipelines for Outcome Prediction From Unstructured Medical Text Notes
IEEE ACCESS
IF3.6
Zero-Shot Medical Image Retrieval for Emerging Infectious Diseases Based on Meta-Transfer Learning - Worldwide, 2020基于元迁移学习的新兴传染病零拍医学图像检索-Worldwide,2020
CHINA CDC WEEKLY
IF2.9

