Return
Artificial intelligence for automated ICD-10 coding: a systematic review of multi-label text classification in clinical narratives
K
S
S
DOI:10.3389/fdgth.2026.1905379.png)
Abstract
En 中文
BackgroundICD-10 coding is an essential process in healthcare systems that supports clinical management; reimbursement; and health data analytics. However; the complexity of its hierarchical structure and the large number of available codes make manual coding limited in terms of time; cost; and consistency. Despite growing research in this area; evidence remains fragmented; particularly regarding real-world implementation readiness.ObjectiveTo review and synthesize existing knowledge on algorithms; datasets; evaluation methods; and real-world implementation readiness of automatic ICD-10 coding systems.MethodsEligible studies were original research articles; preprints; or conference papers published in English between January 1; 2020 and December 31; 2025; and retrieved from seven academic databases: Scopus; PubMed; Web of Science; IEEE Xplore; ACM Digital Library; arXiv; and Google Scholar. Studies were included if they investigated automatic ICD-10 coding from clinical text using machine learning; deep learning; transformer-based; or large language model (LLM) approaches. Methodological quality was assessed using a research-question-driven appraisal framework. This systematic review followed PRISMA 2020 guidance and was preregistered in the Open Science Framework (OSF) at https://osf.io/cegqk.ResultsA total of 257 records were identified; of which 24 studies met the inclusion criteria and contributed 296 experimental evaluations overall. Study quality was high in 7 studies; moderate in 8; and limited by technical or methodological concerns in 9. Hybrid deep learning (Hybrid DL) was most often used as the main automated coding approach; while machine learning (ML) and rule-based approaches were mainly used as baselines. F1-macro was consistently lower than F1-micro among studies reporting both metrics. Hybrid DL showed the most stable performance under all-code or full-code evaluation; while AI model performance varied by the documents-per-label (D/L) ratio.DiscussionThe evidence indicates continued technical progress; particularly through Hybrid DL and transformer-based approaches; while LLM-based methods remain emerging and less consistently effective for structured multi-label coding. The observed D/L–performance relationship suggested that AI model selection should consider dataset structure and label support; in addition to algorithmic complexity.ConclusionAI-based automatic ICD-10 coding is a promising approach for clinical coding support. Future research should prioritize rare-label imbalance; reproducibility; explainability; and validation across diverse clinical settings.Systematic Review Registrationhttps://osf.io/cegqk.
Keywords:
systematic review
deep learning
large language models
explainability
class imbalance
automatic ICD-10 coding
clinical natural language processing
multi-label text classification
Journal
F
IF:
3.8
Papers:
2.0K
Citations:
3.1K
Organization
No organization information available
