arrow
Return

Anatomically guided vision-language model for efficient OCT disease classification and reporting

delete2026-01-01
delete0
PRE
AI
T
Tania Haghighi
S
Sina Gholami
J
Jared T. Sokol
A
Aayush Biswas
J
Jennifer I. Lim
T
Theodore Leng
A
Atalie C. Thompson
H
Hamed Tabkhi
M
Minhaj Nur Alam *
DOI:10.1117/12.3077527delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Optical coherence tomography (OCT) imaging is fundamental for diagnosing retinal diseases. However, existing AI models for OCT lack interpretability, operating as black boxes that limit clinical adoption. To address this, we introduce OCT-BLIP, a compact vision-language model (VLM) that enhances diagnostic performance and provides anatomically grounded, layer-specific explanations. We assembled a multimodal dataset of 40,000 OCT image-text pairs. Images were sourced from both private institutional collections and publicly available repositories. For each image, preliminary layer-specific pathological descriptions spanning five retinal layers were generated using GPT-4o. All generated descriptions underwent subsequent expert review and manual refinement to ensure clinical accuracy and anatomical precision, and consistency. The dataset covers six diagnostic categories: CNV, drusen, DR, GA, DME, and healthy cases. OCT-BLIP, with roughly 247 million parameters based on the BLIP architecture, combines a Vision Transformer encoder and a BERT-style text decoder fine-tuned for classification and captioning. Trained for 50 epochs on 39,000 paired samples using AdamW with a cosine learning-rate schedule, OCT-BLIP achieves 96% classification accuracy, outperforming an 83% standalone ViT baseline and RetinaVLM (< 15%). The model produces precise captions, attaining a mean SBERT similarity of 80.3% and a BERTScore-F1 of 71.5%, substantially surpassing RetinaVLM (71.4% and 42.7%, respectively). The interpretability study assessed structured output clarity, diagnostic usefulness, and anatomical accuracy.
Keywords:
Optical Coherence Tomography
Vision-Language Models
Multimodal Learning
Retinal Disease Classification
Medical Image Captioning

Journal

O
OPHTHALMIC TECHNOLOGIES XXXVI
IF:
0
Papers:
30
Citations:
0

Organization

University of Illinois System cover
University of Illinois System
Scholars:
6.8W
Papers: 6.2W
Citations: 644
U
university of north carolina charlotte
Scholars:
246
Papers: 175
Citations: 0
S
stanford university
Scholars:
1.1W
Papers: 4.2K
Citations: 0
U
University of North Carolina
Scholars:
5.4K
Papers: 2.5K
Citations: 337
researcher View more organizations