Return
OtoVCE: a mechanism-aware language-model evidence layer for hereditary hearing-loss variant interpretation
S
L
P
DOI:10.3389/fgene.2026.1880490.png)
Abstract
En 中文
BackgroundMissense variants in genes implicated in hereditary hearing loss are frequently returned to clinicians as variants of uncertain significance. In silico predictors and rule-based ACMG/AMP frameworks each capture only a fraction of the required evidence; and neither directly accesses case-level and functional evidence reported in the published literature.MethodsWe developed OtoVCE; a four-stage framework that places a large language model inside a calibrated ACMG/AMP rule engine as a structured evidence extractor rather than an end-to-end classifier. Rule-encodable evidence—six in silico missense predictors; gnomAD allele frequencies; UniProt domain annotations; and a hearing-loss-specific protein language model with a pathogenicity head and a mechanism head—is aggregated by the ClinGen Hearing Loss VCEP rule set. For variants that remain uncertain; OtoVCE retrieves the literature from five sources and prompts the language model using the Brnich functional evidence rubric; with the predicted disease mechanism as context; returning PS3; BS3; and PS4 strength assignments with PubMed-verified citations. Rule-based and language-model-derived strengths are then combined into a posterior probability mapped to the five ACMG/AMP classes.ResultsOn a held-out post-2024 ClinVar cohort (N = 1; 885); OtoVCE achieved an area under the receiver-operating curve (AUC) of 0.997 and a sensitivity of 93.6% at a 1% false-positive rate. Performance was preserved across an OTOF gene-leave-out cohort (AUC = 0.993); the external Deafness Variation Database (AUC = 0.917); and the protein-language-model training cohort itself; with model-derived rules disabled (AUC = 0.976); excluding training-set memorization. A paired A/B comparison (N = 1; 930) attributed the literature evidence contribution to the mechanism cue: 7.1-fold more cited identifiers; 11.6-fold more quantitative evidence; and diagnostic strength scores (paired-bootstrap ΔAUC +0.18). OtoVCE identified 211 of 2; 389 uncertain variants as one piece of supporting evidence short of likely pathogenic at 98.1% precision; agreed with ClinGen-VCEP curation at Cohen’s κ = 0.57; and verified 100% of 3; 185 cited PubMed identifiers.ConclusionA literature-grounded; mechanism-aware language-model evidence layer can serve as a complementary component within a calibrated ACMG/AMP workflow for clinical variant interpretation in hereditary hearing loss.
Keywords:
large language model
clinical decision support
hereditary hearing loss
ACMG/AMP variant interpretation
mechanism-aware prompting
Journal
IF:
2.8
Papers:
1.4K
Citations:
4.4W
