Return
Anatomically guided vision-language model for efficient OCT disease classification and reporting
DOI:10.1117/12.3077527.png)
Abstract
En 中文
Optical coherence tomography (OCT) imaging is fundamental for diagnosing retinal diseases. However, existing AI models for OCT lack interpretability, operating as black boxes that limit clinical adoption. To address this, we introduce OCT-BLIP, a compact vision-language model (VLM) that enhances diagnostic performance and provides anatomically grounded, layer-specific explanations. We assembled a multimodal dataset of 40,000 OCT image-text pairs. Images were sourced from both private institutional collections and publicly available repositories. For each image, preliminary layer-specific pathological descriptions spanning five retinal layers were generated using GPT-4o. All generated descriptions underwent subsequent expert review and manual refinement to ensure clinical accuracy and anatomical precision, and consistency. The dataset covers six diagnostic categories: CNV, drusen, DR, GA, DME, and healthy cases. OCT-BLIP, with roughly 247 million parameters based on the BLIP architecture, combines a Vision Transformer encoder and a BERT-style text decoder fine-tuned for classification and captioning. Trained for 50 epochs on 39,000 paired samples using AdamW with a cosine learning-rate schedule, OCT-BLIP achieves 96% classification accuracy, outperforming an 83% standalone ViT baseline and RetinaVLM (< 15%). The model produces precise captions, attaining a mean SBERT similarity of 80.3% and a BERTScore-F1 of 71.5%, substantially surpassing RetinaVLM (71.4% and 42.7%, respectively). The interpretability study assessed structured output clarity, diagnostic usefulness, and anatomical accuracy.
Keywords:
Optical Coherence Tomography
Vision-Language Models
Multimodal Learning
Retinal Disease Classification
Medical Image Captioning
Journal
O
IF:
0
Papers:
30
Citations:
0

